OpenAI says its AI agents log 3.1 hours of runtime for every hour its researchers work

Key points
- OpenAI: 3.1 hours of agent runtime per human work hour across its research org
- Median researcher's daily agent usage exceeds $600 at API prices; top users $7,000
- Over half of successful tasks estimated at 4 to 8 human hours involved intervention
- Every number is OpenAI measuring itself, and it calls them preliminary
OpenAI says its AI agents now put in more hours than its human researchers do. In a post published Sunday, September 6, the company reported that its coding agents log 3.1 hours of runtime for every hour its researchers work. This measures how long agents run, not how much useful research they produce. The company also said it has hit the goal it set last fall of having an "automated research intern" by this September. "By 'research intern,' we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days," the post says.
The figures come from OpenAI's own internal measurements. "Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary," the post says. The post also sets the next target, an "automated AI researcher by March of 2028," describes what the agents still can't do, and explains why OpenAI paused some training after July's Hugging Face incident.
Advertisement
What did OpenAI measure?
OpenAI measured how often researchers use coding agents, tools that write and run code. At the start of the year, OpenAI says, its median researcher used them "only in modest amounts." By mid-August the median researcher was "using more than $600 per day of inference at API prices," and a researcher using more than 90 percent of colleagues was using "more than $7,000 of tokens per day." Inference means running a trained model; the dollar figures value that usage at API prices rather than disclose OpenAI's actual costs. Tokens are chunks of text that AI services count to calculate usage charges. Before June, total agent runtime across the research organization was less than total human labor. By mid-August it was 3.1 times as much.
Researchers also ran more experiments on average during 2026, and August was "an all-time high since tracking began in Jan 2025." OpenAI links that to wider use of Codex, its coding agent, but adds that "our available compute has also grown significantly since 2025," which makes it difficult to separate the effects of agent adoption from increased computing capacity.
OpenAI grouped agents' tasks into six research categories developed by Epoch AI, an outside research group. All six grew between January and August. Writing research and infrastructure code is still the biggest category, with the largest gains in technical help and monitoring training runs. OpenAI says agents still spend very little of their output on high-level planning.
What can the agents not do yet?
Longer tasks still require human oversight. OpenAI says agent success rates rose from January to July across task lengths, but "agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions." People, the post says, "still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems."
The company also warns against reading the activity numbers as progress numbers. "AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won't keep pace with these specific metrics," the company wrote. As agents take over routine tasks, researchers spend more of their time on work that is harder to automate, and computing power, which OpenAI calls compute, "may become more important over time as other bottlenecks diminish."
Internal teams that used to hold office hours to help researchers debug experiments have seen attendance fall this year, and one has stopped holding them. Posts to one of the main internal technical-support channels have declined, and, to OpenAI's knowledge, the questions didn't move to another human-run channel. OpenAI's reading is that researchers are getting that help from agents instead.
Advertisement
What happened after the Hugging Face incident?
The post describes how the July attack changed OpenAI's own work. On July 20, "following the discovery that agents had compromised our research infrastructure," OpenAI temporarily shut down a service used to train models, then restarted it with stricter security controls. Computing power used for reinforcement learning, the training stage where a model learns from feedback on its answers, fell sharply while teams adjusted, and that training paused for two weeks on the newest models meant for release. Most GPU allocation for Astra-class reinforcement-learning experiments between July 20 and August 6, the post says, went to testing the new safety and security controls.
On August 7, early evidence that Astra might have critical cyber capabilities under OpenAI's internal safety rules forced the model into higher-security environments. Over the following week, GPU allocation to Astra fell 59.2 percent within the reinforcement-learning workloads OpenAI studied. Allocation to other models rose 17.2 percent. Because the two groups started at different sizes, that increase offset about 85 percent of Astra's decline. Researchers had shifted much of the computing capacity to other models. OpenAI's reading: "When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise."
Why does this matter for investors?
For stock-market investors, the findings may matter to publicly traded companies that supply OpenAI with chips and computing services. Agent usage valued at more than $600 a day for the median researcher, and $7,000 for one using more than 90 percent of colleagues, is demand for computing that OpenAI creates for itself before any customer asks a question. Rising internal agent usage could add to demand for computing capacity, although the post does not identify the suppliers serving these workloads or quantify additional revenue. For Nvidia (NVDA), Oracle (ORCL) and Microsoft (MSFT), the relevant line in the post is that computing power "may become more important over time as other bottlenecks diminish." On the same Sunday, Nvidia's Jensen Huang said Astra was trained on more than 100,000 of his company's GPUs, with 400,000 more coming.
The post also frames all of this as a step toward what OpenAI calls recursive self-improvement, or RSI, meaning AI that makes the next AI better, and says it doesn't yet know how to do that safely. "We do not yet know how to safely get all the way to aligned, full RSI," the company wrote. OpenAI says it will slow or stop development if it cannot adequately control the risks. OpenAI's next goal is an automated AI researcher by March 2028.
Advertisement