Reasoning's Reckoning: Test-Time Compute and the Price of a Correct Answer
Reasoning models turn accuracy into a line item; the ARC Prize leaderboard, vendor benchmark tables from Anthropic and OpenAI, and the BRAID paper show what a correct answer costs in September 2026.
AI is getting better, faster, stronger and cheaper.
By the numbers
- ARC-AGI-2, GPT-6 Astra (Max)
- 95.0% at $1.12/task
- ARC Prize leaderboard as viewed Sept. 4, 2026; human panel 100% at $17/task · [1] arcprize.org
- ARC-AGI-2, GPT-5.6 Sol, Low to Max reasoning
- 42.5% to 92.5%
- $0.32 to $1.44 per task; same weights, different test-time budget · [1] arcprize.org
- Cheapest ARC-AGI-2 result above 84%
- $0.249/task
- Gemini 3.7 Flash (High), 84.6%; Gemini 3 Deep Think posts the same score at $13.62 · [1] arcprize.org
- ARC-AGI-2, DeepSeek V4 Flash (Max)
- 61.4% at $0.042/task
- Cheapest result in the 60% class on the board · [1] arcprize.org
- BRAID peak performance per dollar
- 74.06×
- gpt-4.1 generator to gpt-5-nano-minimal solver on GSM-Hard, normalized to GPT-5-medium = 1.0 · [9] arXiv (2512.15959)
Claude Opus 5 scores 90.4% on ARC-AGI-2 at $2.06 per task, according to the ARC Prize leaderboard as viewed Sept. 4, 2026, six weeks after Anthropic released the model on July 24, 2026 at $5 per million input tokens and $25 per million output tokens, the day Axios reporter Madison Mills wrote that AI was getting better, faster, stronger and cheaper. Cheaper is the contested word. Reasoning models have converted accuracy into a purchasable quantity, and the same leaderboard prices the human panel’s 100% at $17 per task, GPT-6 Astra’s 95.0% at $1.12, Gemini 3.7 Flash’s 84.6% at $0.249 and DeepSeek V4 Flash’s 61.4% at $0.042. A correct answer now has a bill of materials. This article reads the bill: the test-time compute curve, accuracy per dollar across tiers, the vendor claims of Sept. 1 and July 9, the architectural shift toward single-call agents, and the structured-prompting counterpoint from the BRAID paper.
Reasoning Models as a Purchasable Quantity: The ARC-AGI-2 Cost Curve
The cleanest evidence that reasoning is bought by the unit comes from a single model at two budgets. GPT-5.6 Sol scores 42.5% on ARC-AGI-2 at $0.32 per task with its reasoning effort set to Low and 92.5% at $1.44 with the effort set to Max, according to the ARC Prize leaderboard; the weights are identical, and a 4.5× increase in spend buys a 2.2× increase in accuracy. Claude Opus 4.5 shows the same shape from a lower base, rising from 7.8% with extended thinking switched off to 37.6% with a 64,000-token thinking budget at $2.40 per task. Test-time compute is the variable in both cases. The buyer chooses a point on the curve, and the curve is public.
| Model (setting) | ARC-AGI-2 score | Cost per task | Points per dollar |
|---|---|---|---|
| Human panel | 100% | $17.00 | 5.9 |
| GPT-6 Astra (Max) | 95.0% | $1.12 | 84.8 |
| GPT-5.6 Sol (Max) | 92.5% | $1.44 | 64.2 |
| Claude Opus 5 (Max) | 90.4% | $2.06 | 43.9 |
| Claude Fable 5.1 (Max) | 90.0% | $4.49 | 20.0 |
| Claude Fable 5 (Max) | 89.2% | $5.45 | 16.4 |
| GPT-5.5 (XHigh) | 85.0% | $1.87 | 45.5 |
| Gemini 3.7 Flash (High) | 84.6% | $0.249 | 339.8 |
| Gemini 3 Deep Think (2/26) | 84.6% | $13.62 | 6.2 |
| Gemini 3.1 Pro (Preview) | 77.1% | $0.962 | 80.1 |
| Grok 4.6 (XHigh) | 67.1% | $0.757 | 88.6 |
| DeepSeek V4 Flash 0731 (Max) | 61.4% | $0.042 | 1,461.9 |
| GPT-5.6 Sol (Low) | 42.5% | $0.32 | 132.8 |
Scores and costs are the ARC Prize Foundation’s as viewed Sept. 4, 2026; the points-per-dollar column is this journal’s arithmetic on those figures. Astra’s 95.0% at $1.12 places a frontier model at one-fifteenth of the human panel’s cost, and its ARC-AGI-1 result of 97.5% at $0.433 shows the older benchmark saturating at a price below half a dollar. ARC-AGI-3, the interactive agent benchmark, remains expensive and harness-sensitive: Astra posts 62.7% at $26,100 on the Standard harness and 98.6% at $17,300 on the Provider Adapter harness, Claude Opus 5 (High) posts 30.2% at $20,700 and Sol posts 7.8%, which means the same model’s score can move 36 points on harness choice alone.
Accuracy per Dollar: Where Gemini 3.7 Flash and DeepSeek V4 Flash Break the Frontier
Ranking by points per dollar inverts the leaderboard. DeepSeek V4 Flash delivers 61.4% at $0.042, roughly 1,462 points per dollar and 17 times Astra’s ratio, on a model DeepSeek previewed on April 24, 2026 after training it partly on Huawei Ascend chips, according to Reuters. Gemini 3.7 Flash delivers 84.6% at $0.249, which beats Gemini 3.1 Pro’s 77.1% at $0.962 on both axes and matches Gemini 3 Deep Think’s 84.6% at one fifty-fifth of Deep Think’s $13.62. Inside Anthropic’s lineup the marginal point costs the most. Opus 5 at $2.06 scores 90.4%; Fable 5.1 at $4.49 scores 90.0%; Fable 5 at $5.45 scores 89.2%. The frontier tier from every vendor pays a steep premium for the last few points, and the premium is the price of certainty on the hardest 5% of tasks. Where a task distribution resembles ARC-AGI-2’s, a buyer who routes 90% of traffic to a Flash-class model and escalates the remainder to a Max-tier model spends a fraction of the all-frontier budget. That arithmetic is the business case for tiered routing, and it is why the “best LLM for agents” question has become a portfolio question.
Vendor Claims: Fable 5.1, GPT-5.6 Sol and the Token-Efficiency Contest
Anthropic published a benchmark table with the Sept. 1, 2026 release of Claude Fable 5.1 and Mythos 5.1, and the table is vendor-published. Fable 5.1 scores 55.8% on Terminal-Bench 4.0 against 42.0% for Fable 5, 52.3% for Opus 5 and 37.3% for GPT-5.6 Sol; 52.6% on Terminal-Bench-Science 0.1 against 24.7%, 29.0% and 22.4%; 60.9% on Humanity’s Last Exam with tools disabled against 57.8% and 56.6%; 77.9% on OSWorld 2.0 (partial) against 72.9% and 75.4%; and 73.4% on CursorBench 3.2.0 against 70.5%, 70.0% and 67.2%. Pricing stayed at $10 per million input tokens and $50 per million output tokens, cache reads fell 75% to $0.25 per million, and Anthropic claims cost reductions of around 25% on typical workloads and up to around 45% on highly agentic workloads relative to Fable 5.
OpenAI made the mirror-image claim on July 9, 2026. CEO Sam Altman said GPT-5.6 Sol is “54% more token efficient” for coding, TechCrunch reported at launch, and the company said Sol uses under half the output tokens of competitors at about one-third lower cost; Sol scored 80 on the Artificial Analysis Coding Agent Index, 2.8 points above Fable 5. Sol’s list price is $5 per million input tokens and $30 per million output tokens, which Value Add VC described on Aug. 27, 2026 as OpenAI’s first flagship price increase since GPT-4 in March 2023, four times the $1.25 input price of the legacy GPT-5. Both vendors now sell efficiency as a headline feature. Each claim still awaits independent replication, and both should be read as marketing until one appears.
| Benchmark (vendor-published, Sept. 1, 2026) | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 55.8% | 42.0% | 52.3% | 37.3% |
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Humanity’s Last Exam, tools disabled | 60.9% | 57.8% | 56.6% | — |
| OSWorld 2.0 (partial) | 77.9% | 72.9% | 75.4% | — |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Single-Call Solutions: How Reasoning Models Rewired the Agent Stack
Paolo Perrone, writing “The AI Agents Stack (2026 Edition)” for O’Reilly Radar on June 8, 2026, observed that reasoning models moved agents “from multistep chains to single-call solutions.” That shift moved cost from the orchestration layer into the model layer: where a 2024 agent decomposed a task across a dozen cheap calls, a 2026 agent hands the whole task to one reasoning call with a large thinking budget. Menlo Ventures’ Dec. 9, 2025 survey recorded the market share that followed, with Anthropic at 40% of enterprise LLM API spend, up from 24%, OpenAI at 27%, down from 50% in 2023, and Google at 21%, up from 7%; in coding, Anthropic held 54%. The backdrop is the long price decline the Stanford AI Index documented in April 2025: the cost of querying a GPT-3.5-level model fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a 280× drop in about 18 months. Test-time compute reverses part of that decline at the task level, because a single-call agent consumes thinking tokens in proportion to difficulty, and difficulty is decided at run time. Latency follows the same curve. A Max-tier answer arrives later than a Low-tier answer, and an agent with a fixed latency budget therefore buys accuracy with both dollars and seconds.
The Structured-Prompting Counterpoint: BRAID’s 74× Performance per Dollar
Armağan Amcalar and Eyup Cinar of OpenServ Labs posted “BRAID: Bounded Reasoning for Autonomous Inference and Decisions” to arXiv on Dec. 17, 2025, and the paper argues that structure can substitute for budget. A capable generator model writes a bounded reasoning graph in Mermaid syntax, with computed values masked; a smaller solver model follows the graph as its system prompt; and a GPT-5.2 adjudicator scores free-form answers. Performance per dollar is normalized to GPT-5-medium at 1.0. The headline result pairs a gpt-4.1 generator with a gpt-5-nano-minimal solver on GSM-Hard for 96% accuracy at a performance-per-dollar ratio of 74.06; on SCALE MultiChallenge gpt-4o rises from 19.9% to 53.7% and on AdvancedIF gpt-5-nano-minimal rises from 18% to 40%. Amcalar, chief technology officer of OpenServ Labs, said in the March 6, 2026 press release announcing the results that the method lets a team “run 30 different solution paths for the price of one.” The authors acknowledge that the graphs are LLM-generated and static, that GSM-Hard sits near its ceiling with baselines above 90% and possible contamination, and that they judged answers with an LLM because forced output schemas degrade generation. CryptoSlate’s Liam Wright wrote on April 6, 2026 that the benchmark methodology and task selection await independent replication, and the 472-question, roughly 100,000-run evaluation described in the release remains vendor-adjacent evidence until a third party reproduces it. Read against the ARC curve, BRAID names a second lever. Budget buys accuracy along one axis; structure, if the parity effect holds beyond the OpenAI model family, buys it along another.
What to Watch
Four measurements will settle how much a correct answer costs by early 2027. The ARC Prize Foundation’s next ARC-AGI-3 harness update will show whether Astra’s 36-point gap between Standard and Provider Adapter runs narrows, because harness effects of that size make cost-per-task comparisons between vendors provisional. Independent replications of Anthropic’s 25% to 45% cost-reduction claims and OpenAI’s 54% token-efficiency claim will convert marketing into data, and Artificial Analysis’s index is the most likely venue. DeepSeek’s V4 pricing, which the company has yet to disclose alongside its Ascend training claim, will determine whether $0.042 per task at 61.4% is a durable price or a subsidized one. The third-party replication that CryptoSlate called for on April 6 would establish whether bounded reasoning generalizes beyond the OpenAI models BRAID tested, and with it whether structured prompting belongs in the routing layer of every agent stack or in a footnote.
Sources
14 cited · AP style
- ARC Prize Foundation, “ARC Prize Leaderboard”, arcprize.org, As viewed Sept. 4, 2026. arcprize.org
- Madison Mills, “Anthropic Releases New Model, Claude Opus 5”, Axios, July 24, 2026. axios.com
- Anthropic, “Claude Fable 5.1 and Mythos 5.1”, Anthropic, Sept. 1, 2026. anthropic.com
- TechCrunch, “OpenAI Launches Its New Family of Models With GPT-5.6”, TechCrunch, July 9, 2026. techcrunch.com
- OpenAI, “API Pricing”, OpenAI, As viewed Sept. 4, 2026. openai.com
- Anthropic, “Pricing”, Claude Developer Platform, As viewed Sept. 4, 2026. platform.claude.com
- Menlo Ventures, “2025: The State of Generative AI in the Enterprise”, Menlo Ventures, Dec. 9, 2025. menlovc.com
- Paolo Perrone, “The AI Agents Stack (2026 Edition)”, O'Reilly Radar, June 8, 2026. oreilly.com
- Armağan Amcalar and Eyup Cinar, “BRAID: Bounded Reasoning for Autonomous Inference and Decisions”, arXiv (2512.15959), Dec. 17, 2025. arxiv.org
- Coyotiv and OpenServ Labs, “Coyotiv and OpenServ Labs Demonstrate Up to 74x AI Reasoning Efficiency Gains in New Research”, Newsfile, March 6, 2026. newsfilecorp.com
- Liam Wright, “Crypto AI Project OpenServ Says It Can Beat OpenAI, but the Real Test Starts Now”, CryptoSlate, April 6, 2026. cryptoslate.com
- Stanford HAI, “AI Index 2025: State of AI in 10 Charts”, Stanford Institute for Human-Centered AI, April 2025. hai.stanford.edu
- Reuters, “China's AI Darling DeepSeek Previews New Model”, Reuters via Investing.com, April 24, 2026. investing.com
- Value Add VC, “OpenAI API Pricing 2026: GPT-4o, o3 and GPT-5 Cost Breakdown for Developers”, Value Add VC, Aug. 27, 2026. valueaddvc.com
Related reading
Bounded Brilliance: How BRAID Bends the Cost Curve of Machine Reasoning
A December 2025 arXiv paper from OpenServ Labs argues that bounded reasoning graphs, written in Mermaid and handed to nano-class models, deliver flagship accuracy at a fraction of the price.
8 min · 14 sources
Tokens, Tallied: The Economics of Inference for Agentic Workloads
Token economics for agentic workloads pit a 280× collapse in LLM inference cost against quadrillion-token volumes; here are the prices, the spending forecasts and the FinOps levers, dated to September 2026.
7 min · 14 sources
Benchmarks, Broken and Better: Evaluating Agents from SWE-bench to ARC-AGI-3
AI agent benchmarks now carry cost per task, confidence intervals and contamination warnings, and the careful buyer treats every leaderboard as an instrument with a stated error bar.
8 min · 14 sources