OpenAI's GPT-6 Astra beats Claude Fable 5.1 at 39% lower cost per correct task in Signal65 test

Key points
- Signal65: 279 of 280 enterprise jobs finished, zero fabricated answers
- $1.51 per correct task against $2.46 for Claude Fable 5.1
- Code Arena ranks GPT-6 Astra (Max) first at 1,797, on 1,199 votes
- Neither test covers Astra's headline computer-use claims
OpenAI's new model, GPT-6 Astra, has now been tested by two groups outside OpenAI, and both put it ahead of Anthropic's Claude Fable 5.1. The testing firm Signal65 said Astra "completed 279 of 280 real multi-step enterprise jobs end to end and fabricated nothing on the unanswerable questions," and did it for $1.51 per correct task against $2.46 for Claude. On Arena's Code Arena, where people vote on which model builds the better web page, Astra moved into first place.
Both results have limits, and this piece goes through them, along with the demos people are posting of what Astra can do and the public companies that carry OpenAI's models to customers. Signal65 posted its numbers on X on Saturday, September 5, and the Arena leaderboard reflects votes through the same day.
When OpenAI launched Astra on September 3, every benchmark number came from OpenAI. President Greg Brockman closed that briefing with "welcome to the AGI era." The model went first to enterprise customers in the Daybreak program, then to paid ChatGPT users and the API, at $10 per million input tokens and $50 per million output tokens. Anthropic charges the same list price for Claude Fable 5.1.
Advertisement
What does the Arena ranking measure?
Code Arena shows people two anonymous models building the same web page, and they vote for the one they prefer. As of September 5 the WebDev board had 650,961 votes across 126 models. GPT-6 Astra (Max) sat first at 1,797 and Claude Fable 5.1 (Max) second at 1,762. Claude Opus 5 (Max) was third at 1,688, one point ahead of Alibaba's Qwen 3.8 Max 0902.
The vote counts are lopsided. Astra's score rests on 1,199 votes, with a confidence interval of plus or minus 24 points. Fable's rests on 2,275 votes, plus or minus 16. Arena assigns the two models separate rank positions, first and second, rather than a shared range. That call rests on Arena's own statistical method, and the two intervals do overlap slightly at the edges. The board measures preference on web development tasks only. It says nothing about spreadsheets, research or the computer-use work OpenAI led with at launch.
What did Signal65 test?
Signal65's PINNACLE benchmark runs each model through 280 enterprise jobs, 160 on governed company data and 120 on messy "as-found" data, with more than 80 rounds of tool use in some jobs, according to the firm's methodology. A job counts only if it comes out right end to end. A separate retrieval suite mixes in questions the supplied documents cannot answer, to measure how often a model makes something up. Grading is done in code against answer keys held outside the agent's sandbox, with no human raters and no model grading another model.
On September 5 Signal65 posted Astra's results. At maximum reasoning effort the model finished 279 of the 280 jobs and fabricated nothing on the unanswerable questions. Signal65 has not published how many unanswerable questions the suite contains, so that zero is a rate without a stated denominator. At medium effort it finished all 280 but invented answers to 2.3 percent of the unanswerable questions. Claude Fable 5.1, tested on September 2, finished 276 of 280 and fabricated on 0.7 percent. Signal65 said Astra at maximum effort produced 51 percent fewer weighted errors than Fable, the previous leader on its table, and 55 percent fewer than Meta's Muse Spark 1.3.
The cost figure uses list prices for input, cached input and output tokens, and charges the tokens spent on failed attempts to the tasks that finished. Since Astra and Fable list at the same $10 and $50, the difference is how many tokens each one burns. Signal65 put Astra at $1.51 per correct task at maximum effort and $0.94 at medium, against $2.46 for Fable. The firm said medium effort wrote 2.5 times fewer output tokens per correct task than maximum. "The extra 57 cents per correct task buys the refusal behavior rather than the job completion, so the effort setting is a deployment decision," Signal65 wrote.
Signal65 describes itself as an independent analyst and testing firm. Its governance page says it chose the subject, funded the PINNACLE work and publishes regardless of outcome, and that vendors get a factual-accuracy review before publication but no approval right and no veto. Nvidia and AMD supplied hardware and engineering support for the benchmark, with the same review window as everyone else. Astra and Fable were scored over their makers' hosted APIs. PINNACLE lists Kamiwaza.ai as a partner.
Advertisement
What is Astra doing outside the benchmarks?
The examples circulating are demos, not audits. Yunfan Ye, an AI researcher who worked at Google and Meta and now runs Rome AI Lab, posted on September 3 that he gave Astra a Zillow listing and it built a 3D model of the house from the listing photos and produced a promotional video. "The video was just created in one shot and there are still some wrong details," he wrote.
Anshu, a former Y Combinator founder who worked on UI/UX and AI at Apple, posted on September 4 that Astra built a 3D game in one shot in 45 minutes, using "hardly a couple % of my quota." The trick to good graphics, Anshu wrote, was having the model generate the images itself.
OpenAI's own launch post read, "Anything you can do on a computer, Astra can do for you. Fast." At the briefing Brockman said the model "can zip through spreadsheets, fill out forms, and navigate across web pages often at superhuman speed," Fortune reported. Neither Arena nor Signal65 tested that computer-use claim. OpenAI's own figure was 72.6 percent on the OSWorld 2.0 computer-use test, and that number is still the company's.
What does it mean for the public companies around OpenAI?
OpenAI and Anthropic are both private. Astra reaches customers through OpenAI's API and through Microsoft (MSFT) Azure and Amazon Web Services, and it was pretrained at the Stargate site in Texas that Oracle (ORCL) builds and runs. Anthropic's revenue passes through some of the same channels, since Claude is sold on both Amazon Bedrock and Google Cloud, and Meta has said it could spend $10 billion a year on Anthropic's models.
Signal65 put the trade-off between Astra's two settings in one sentence. "An agent answering customer or policy questions with no human in the path wants maximum, and an agent working inside a workflow that ends with a reviewer can run at medium and keep the savings."
Advertisement