Advertisement

OpenAI says its AI agents log 3.1 hours of runtime for every hour its researchers work

OpenAI logo on a black background

Key points

  • OpenAI: 3.1 hours of agent runtime per human work hour across its research org
  • Median researcher's daily agent usage exceeds $600 at API prices; top users $7,000
  • Over half of successful tasks estimated at 4 to 8 human hours involved intervention
  • Every number is OpenAI measuring itself, and it calls them preliminary

OpenAI says its AI agents now put in more hours than its human researchers do. In a post published Sunday, September 6, the company reported that its coding agents log 3.1 hours of runtime for every hour its researchers work. This measures how long agents run, not how much useful research they produce. The company also said it has hit the goal it set last fall of having an "automated research intern" by this September. "By 'research intern,' we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days," the post says.

The figures come from OpenAI's own internal measurements. "Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary," the post says. The post also sets the next target, an "automated AI researcher by March of 2028," describes what the agents still can't do, and explains why OpenAI paused some training after July's Hugging Face incident.

Advertisement

What did OpenAI measure?

OpenAI measured how often researchers use coding agents, tools that write and run code. At the start of the year, OpenAI says, its median researcher used them "only in modest amounts." By mid-August the median researcher was "using more than $600 per day of inference at API prices," and a researcher using more than 90 percent of colleagues was using "more than $7,000 of tokens per day." Inference means running a trained model; the dollar figures value that usage at API prices rather than disclose OpenAI's actual costs. Tokens are chunks of text that AI services count to calculate usage charges. Before June, total agent runtime across the research organization was less than total human labor. By mid-August it was 3.1 times as much.

Researchers also ran more experiments on average during 2026, and August was "an all-time high since tracking began in Jan 2025." OpenAI links that to wider use of Codex, its coding agent, but adds that "our available compute has also grown significantly since 2025," which makes it difficult to separate the effects of agent adoption from increased computing capacity.

OpenAI grouped agents' tasks into six research categories developed by Epoch AI, an outside research group. All six grew between January and August. Writing research and infrastructure code is still the biggest category, with the largest gains in technical help and monitoring training runs. OpenAI says agents still spend very little of their output on high-level planning.

What can the agents not do yet?

Longer tasks still require human oversight. OpenAI says agent success rates rose from January to July across task lengths, but "agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions." People, the post says, "still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems."

The company also warns against reading the activity numbers as progress numbers. "AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won't keep pace with these specific metrics," the company wrote. As agents take over routine tasks, researchers spend more of their time on work that is harder to automate, and computing power, which OpenAI calls compute, "may become more important over time as other bottlenecks diminish."

Internal teams that used to hold office hours to help researchers debug experiments have seen attendance fall this year, and one has stopped holding them. Posts to one of the main internal technical-support channels have declined, and, to OpenAI's knowledge, the questions didn't move to another human-run channel. OpenAI's reading is that researchers are getting that help from agents instead.

Advertisement

What happened after the Hugging Face incident?

The post describes how the July attack changed OpenAI's own work. On July 20, "following the discovery that agents had compromised our research infrastructure," OpenAI temporarily shut down a service used to train models, then restarted it with stricter security controls. Computing power used for reinforcement learning, the training stage where a model learns from feedback on its answers, fell sharply while teams adjusted, and that training paused for two weeks on the newest models meant for release. Most GPU allocation for Astra-class reinforcement-learning experiments between July 20 and August 6, the post says, went to testing the new safety and security controls.

On August 7, early evidence that Astra might have critical cyber capabilities under OpenAI's internal safety rules forced the model into higher-security environments. Over the following week, GPU allocation to Astra fell 59.2 percent within the reinforcement-learning workloads OpenAI studied. Allocation to other models rose 17.2 percent. Because the two groups started at different sizes, that increase offset about 85 percent of Astra's decline. Researchers had shifted much of the computing capacity to other models. OpenAI's reading: "When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise."

Why does this matter for investors?

For stock-market investors, the findings may matter to publicly traded companies that supply OpenAI with chips and computing services. Agent usage valued at more than $600 a day for the median researcher, and $7,000 for one using more than 90 percent of colleagues, is demand for computing that OpenAI creates for itself before any customer asks a question. Rising internal agent usage could add to demand for computing capacity, although the post does not identify the suppliers serving these workloads or quantify additional revenue. For Nvidia (NVDA), Oracle (ORCL) and Microsoft (MSFT), the relevant line in the post is that computing power "may become more important over time as other bottlenecks diminish." On the same Sunday, Nvidia's Jensen Huang said Astra was trained on more than 100,000 of his company's GPUs, with 400,000 more coming.

The post also frames all of this as a step toward what OpenAI calls recursive self-improvement, or RSI, meaning AI that makes the next AI better, and says it doesn't yet know how to do that safely. "We do not yet know how to safely get all the way to aligned, full RSI," the company wrote. OpenAI says it will slow or stop development if it cannot adequately control the risks. OpenAI's next goal is an automated AI researcher by March 2028.

Advertisement

Frequently asked questions

What did OpenAI announce on September 6, 2026?

In a post titled Research acceleration: The view inside OpenAI, the company said it has reached its goal of an automated research intern, a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. It said its research organization now logs 3.1 agent-workdays of runtime for every eight-hour workday of human labor as of mid-August 2026, a measure of agent running time rather than of research output, and that it aims to build an automated AI researcher by March 2028.

How much does OpenAI spend on AI agents per researcher?

OpenAI said that by mid-August 2026 its median researcher was using more than $600 per day of inference at API prices, and the 90th percentile user in its research organization more than $7,000 of tokens per day. At the start of the year the median researcher used coding agents only in modest amounts.

Can OpenAI's agents do research on their own?

Not yet, by OpenAI's account. It said agents still require significant human steering, especially as task complexity rises, and that in the six months to July 2026 more than half of successful 4-to-8-hour tasks involved one or more human interventions. People still set research priorities and decide whether to scale, pause or deploy systems.

Are OpenAI's research acceleration numbers independently verified?

No. The figures come from OpenAI's internal measurements of its own research organization, and the company describes its measurement efforts as preliminary. It also cautioned that the overall pace of research progress likely will not keep pace with these activity metrics, because the least automatable tasks become the bottleneck as automation advances.

More on MSFT and NVDA

Dennis Singleton
Dennis Singleton

Dennis Singleton has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.

OpenAI: 3.1 hours of agent runtime per human hour in its labs