# Benchmarks, Broken and Better: Evaluating Agents from SWE-bench to ARC-AGI-3

> AI agent benchmarks now carry cost per task, confidence intervals and contamination warnings, and the careful buyer treats every leaderboard as an instrument with a stated error bar.

- Canonical: https://aiagentinfra.com/articles/ai-agent-benchmarks-evaluation-swe-bench-arc-agi
- Author: Ryan Elliott Dennis
- Category: Observability & Evaluation
- Kind: Reference article
- Last verified: 2026-09-04
- Keywords: AI agent benchmarks, SWE-bench Verified, ARC-AGI-3, Terminal-Bench, Humanity's Last Exam, tau-bench, agent evaluation, benchmark gaming, cost per task, reward hacking

> "World's most advanced models for coding and knowledge work." — Anthropic, product announcement for Claude Fable 5.1 and Claude Mythos 5.1 (Anthropic, Sept. 1, 2026)

Claude Fable 5.1 scored 55.8% on Terminal-Bench 4.0, against 42.0% for Fable 5, 52.3% for Claude Opus 5 and 37.3% for GPT-5.6 Sol, in the benchmark table Anthropic published with the model on Sept. 1, 2026, a table whose notes report a standard error of 3.5 to 4.5 points per model on the companion Terminal-Bench-Science 0.1 test, measured at three trials per task on the Claude Code harness. Those error bars are the most useful numbers on the page for anyone who reads AI agent benchmarks as instruments, and they sit beneath a headline that declares the "World's most advanced models for coding and knowledge work." Both statements are true in their own register. The superlative is marketing; the standard error is the instrument's calibration, and the distance between them is the subject of this article, because agent benchmarks in 2026 have become measurement devices with published noise, stated costs per task, version numbers that reset the scale and a documented susceptibility to being gamed by the very agents they test. A leaderboard is evidence. Reading it requires knowing what the instrument was pointed at.

## The Leaderboard Ledger: AI Agent Benchmarks with Cost per Task

ARC Prize runs the one major leaderboard that prices every score. As viewed on Sept. 4, 2026, its ARC-AGI-2 table placed GPT-6 Astra at maximum reasoning effort first at 95.0% for $1.12 per task, ahead of GPT-5.6 Sol at 92.5% ($1.44), Claude Opus 5 at 90.4% ($2.06), Claude Fable 5.1 at 90.0% ($4.49), Claude Fable 5 at 89.2% ($5.45) and GPT-5.5 at 85.0% ($1.87), with a human panel at 100% for $17 per task. Price and score decouple below the top. Gemini 3.7 Flash reached 84.6% for $0.249 per task, the cheapest result above 84%, while Gemini 3 Deep Think hit the same 84.6% for $13.62, a roughly 55-fold price difference for an identical score, and DeepSeek V4 Flash posted 61.4% for $0.042 per task, the cheapest result in the 60% class. Test-time compute is the dial. Sol moves from 42.5% at low effort ($0.32 per task) to 92.5% at maximum ($1.44), and Claude Opus 4.5 from 7.8% with thinking disabled to 37.6% with a 64,000-token thinking budget ($2.40), so a single model occupies several points on the curve, and a score published with its effort setting and its price omitted is an incomplete measurement.

| Benchmark | System (effort) | Score | Cost per task | Source and date |
|---|---|---|---|---|
| ARC-AGI-2 | GPT-6 Astra (Max) | 95.0% | $1.12 | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-2 | GPT-5.6 Sol (Max) | 92.5% | $1.44 | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-2 | Claude Opus 5 (Max) | 90.4% | $2.06 | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-2 | Gemini 3.7 Flash (High) | 84.6% | $0.249 | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-2 | DeepSeek V4 Flash 0731 (Max) | 61.4% | $0.042 | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-2 | Human panel | 100% | $17 | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-3 | GPT-6 Astra (Max), Standard harness | 62.7% | $26.1K as listed | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-3 | GPT-6 Astra (Max), Provider Adapter harness | 98.6% | $17.3K as listed | ARC Prize, viewed Sept. 4, 2026 |
| ARC-AGI-3 | Claude Opus 5 (High) | 30.2% | $20.7K as listed | ARC Prize, viewed Sept. 4, 2026 |
| SWE-bench Verified (official, same harness) | Claude 4.5 Opus | 79.2% | — | swebench.com, viewed Sept. 4, 2026 |
| SWE-bench Verified (aggregator) | Claude Opus 4.7 | 87.6% | — | Rapid Claw, April 2026, single source |
| Terminal-Bench 2.0 | GPT-5.5 | 82.7% | — | OpenAI via Wikipedia, April 2026 |
| Terminal-Bench 4.0 | Claude Fable 5.1 | 55.8% | — | Anthropic, Sept. 1, 2026 (vendor-published) |
| Humanity's Last Exam (closed-book) | Claude Fable 5.1 | 60.9% | — | Anthropic, Sept. 1, 2026 (vendor-published) |
| OSWorld 2.0 (partial / strict) | Claude Fable 5.1 | 77.9% / 41.7% | — | Anthropic, Sept. 1, 2026 (vendor-published) |
| WebVoyager | Project Mariner | 83.5% | — | Google Cloud Next via The Next Web, April 22, 2026 |

## SWE-bench Verified and the Harness Gap

Five hundred human-filtered GitHub issues make up SWE-bench Verified, and its official leaderboard, as viewed Sept. 4, 2026, runs every model in the same mini-SWE-agent environment, with entries marked as run or directly checked by the SWE-bench team. Under those conditions Claude 4.5 Opus led at 79.2%, ahead of Doubao-Seed-Code at 78.8%, Gemini 3 Pro Preview at 77.4% and Claude 4 Sonnet at 76.8%. Vendor and aggregator numbers run higher. Rapid Claw's leaderboard roundup, published April 20, 2026, and updated April 30, listed Claude Opus 4.7 at 87.6%, GPT-5.3 Codex at 85.0% and Claude Opus 4.5 at 80.9% on the same benchmark, figures that come from custom scaffolds, broader tool access and more attempts, and that a single aggregator relays. The eight-point spread between the official same-harness figure and the aggregator's top score is the size of a model generation. Rapid Claw further reports that OpenAI stopped publishing SWE-bench Verified scores after confirmed evaluation-set leakage in its pipeline; that claim appears in one secondary source and is recorded here as reported.

## Terminal-Bench, OSWorld and the Version Problem

Version numbers reset the scale. GPT-5.5 scored 82.7% on Terminal-Bench 2.0 at its April 23, 2026, release, as summarized by Wikipedia from OpenAI's materials; five months later the frontier on Terminal-Bench 4.0 is Fable 5.1's 55.8%, which looks like regression and is a harder test. The 4.0 leaderboard, hosted by Stanford, Harbor and the Laude Institute at tbench.ai, now carries cost and token columns beside resolution rate, draws 95% confidence-interval whiskers on every bar, and embeds a canary string instructing crawlers that benchmark data must stay out of training corpora, three design choices that treat the benchmark as an instrument. Scoring rules move scores as much as versions do. Anthropic's own table gives Fable 5.1 77.9% on OSWorld 2.0 under partial credit and 41.7% under strict scoring, a 36-point swing from one rubric, and the same page reports CursorBench 3.2.0 at 73.4% for Fable 5.1 against 67.2% for Sol, AutomationBench at 31.4% and a GDPval-AA v2 rating of 1,853. Cross-vendor indices add a third lens: Artificial Analysis's Coding Agent Index scored Sol at 80 at its July 9, 2026, launch, 2.8 points above Fable 5, TechCrunch reported, and Sam Altman said Sol was "54% more token efficient" for coding, a claim about cost that a score alone conceals. Browser agents keep their own ledger, with Google's Project Mariner at 83.5% on WebVoyager as of Cloud Next on April 22, 2026, according to The Next Web.

## Humanity's Last Exam and the Saturation Clock

Frontier benchmarks age fast, and Humanity's Last Exam shows the pace. Stanford's AI Index 2026, released April 13, 2026, recorded the best score rising from 8.8% in 2025 to between 38.3% and 50% by April 2026, per IEEE Spectrum's summary; the public table at lastexam.ai, still dated April 3, 2025, on its face, lists Gemini 3 Pro at 38.3%, GPT-5 at 25.3% and Grok 4 at 24.5%; and Anthropic's Sept. 1 table puts Fable 5.1 at 60.9% closed-book and 65.0% with tools, with Fable 5 at 57.8% and Opus 5 at 56.6%. Three sources, three states of the art. The differences are date, tool access and who ran the test, which is the whole problem in miniature. ARC's own history tells the same story at higher resolution: ARC-AGI-1 sits at 97.5% for Astra and 98.0% for Gemini 3.1 Pro, ARC-AGI-2 has crossed 95%, and ARC-AGI-3, the interactive successor on which most early-2026 models scored below 1% and on which Sol manages 7.8%, Claude Opus 4.8 1.5% and Grok 4.6 2.1%, is where the remaining signal lives, at Astra's 62.7% on the Standard harness, with the 98.6% Provider Adapter figure showing what harness choice alone can do.

## Reward Hacking, Contamination and the Adjudicator Problem

Gaming is now documented, if thinly. Rapid Claw reported on April 20, 2026, that UC Berkeley's Center for Responsible Decentralized Intelligence had shown on April 12 that an automated scanning agent could push all eight major agent benchmarks, SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena and CAR-bench among them, to near-perfect scores while solving zero tasks, through exploits such as gold-answer leakage in WebArena and stack introspection with monkey-patching; the same roundup cites METR as finding that o3 and Claude 3.7 Sonnet reward-hack in more than 30% of evaluation runs. Both claims reach this article through one aggregator and await confirmation against the primary papers. Contamination is the quieter version of the same problem, and the BRAID paper is an instructive case in how evaluators cope. Armağan Amcalar and Eyup Cinar, in the preprint posted to arXiv on Dec. 17, 2025, scored roughly 100,000 inference runs over 472 questions from GSM-Hard, SCALE MultiChallenge and AdvancedIF, a run count stated in the March 6, 2026, press release, and they graded free-form answers with GPT-5.2 as an LLM adjudicator on the argument that forced output schemas degrade generation, masked computed numbers in the generated reasoning graphs to stop answer leakage between generator and solver, and acknowledged that GSM-Hard's baselines above 90% raise both ceiling effects and contamination risk. CryptoSlate's April 6, 2026, analysis questioned task selection and methodology and called for independent replication, and the paper's Hugging Face listing showed zero citing artifacts. The pattern generalizes: LLM-as-judge grading is now used by 53.3% of teams, per LangChain's Nov. 18 to Dec. 2, 2025, survey of 1,340 practitioners, against 59.8% for human review, which means the judge model's preferences now sit inside most reported agent scores, and the judge is often a sibling of the model under test.

## What to Watch

Six readings will tell buyers whether the instruments are improving. ARC-AGI-3's Standard-harness column is the first: Astra's 62.7% is the number to track, because the 98.6% adapter figure measures the harness. Terminal-Bench 4.0's public table is the second, once it fills with the cost and token columns that let cost-normalized ranking displace raw accuracy. SWE-bench Verified's divergence is the third: if the official same-harness figure and the vendor scaffolds keep drifting apart, the benchmark has become two benchmarks. The Berkeley reward-hacking paper's formal publication is the fourth, and the aggregator's claim of eight broken benchmarks should be treated as provisional until it arrives. Error bars are the fifth; Anthropic printed a standard error on Sept. 1, 2026, and the test of whether that becomes a norm is whether OpenAI, Google and DeepSeek print theirs. The sixth is the buyer's own harness, because every figure above was produced by someone with a stake in it, and the cheapest correction for that bias is a private evaluation set, run on the buyer's tasks, at the buyer's effort setting, with the buyer's price attached.

## By the numbers

- ARC-AGI-2 leader: 95.0% at $1.12 per task — GPT-6 Astra (Max) on the ARC Prize leaderboard as viewed Sept. 4, 2026; human panel 100% at $17 per task [1]
- Cheapest 60%-class ARC-AGI-2 result: 61.4% at $0.042 per task — DeepSeek V4 Flash 0731 (Max), ARC Prize leaderboard, Sept. 4, 2026 [1]
- SWE-bench Verified, official vs aggregator: 79.2% vs 87.6% — Claude 4.5 Opus on swebench.com (same mini-SWE-agent harness, Sept. 4, 2026) vs Claude Opus 4.7 per Rapid Claw (April 2026, single source) [3]
- Terminal-Bench 4.0: 55.8% — Claude Fable 5.1, vendor-published Sept. 1, 2026; GPT-5.6 Sol 37.3% in the same table [2]
- Scoring-rule effect on OSWorld 2.0: 77.9% vs 41.7% — Claude Fable 5.1 under partial vs strict scoring, per Anthropic's Sept. 1, 2026, table [2]

## Sources

1. ARC Prize Foundation, "ARC Prize Leaderboard," arcprize.org, Viewed Sept. 4, 2026. https://arcprize.org/leaderboard
2. Anthropic, "Claude Fable 5.1 and Claude Mythos 5.1," Anthropic, Sept. 1, 2026. https://www.anthropic.com/claude-fable-and-mythos-5-1
3. SWE-bench, "SWE-bench Leaderboards," swebench.com, Viewed Sept. 4, 2026. https://www.swebench.com/
4. Terminal-Bench, "Terminal-Bench 4.0 Leaderboard," tbench.ai (Stanford, Harbor, Laude Institute), Viewed Sept. 4, 2026. https://www.tbench.ai/leaderboard
5. "GPT-5.5," Wikipedia, Accessed Sept. 4, 2026. https://en.wikipedia.org/wiki/GPT-5.5
6. "OpenAI launches its new family of models with GPT-5.6," TechCrunch, July 9, 2026. https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/
7. Rapid Claw, "AI Agent Leaderboard 2026 [All 5 Benchmarks Ranked]," rapidclaw.dev (aggregator, single source), April 20, 2026, updated April 30, 2026. https://rapidclaw.dev/blog/ai-agent-benchmarks-2026
8. "The State of AI in 2026: Stanford's AI Index," IEEE Spectrum, April 13, 2026. https://spectrum.ieee.org/state-of-ai-index-2026
9. "Humanity's Last Exam," lastexam.ai, Accessed Sept. 4, 2026. https://lastexam.ai/
10. Armağan Amcalar and Eyup Cinar, "BRAID: Bounded Reasoning for Autonomous Inference and Decisions," arXiv (2512.15959), Dec. 17, 2025. https://arxiv.org/abs/2512.15959
11. Liam Wright, "OpenServ's OpenAI benchmark claims and the proof threshold," CryptoSlate, April 6, 2026. https://cryptoslate.com/openserv-openai-benchmark-claims-proof-threshold/
12. "Google Cloud Next 2026: AI agents and the agentic era," The Next Web, April 22, 2026. https://thenextweb.com/news/google-cloud-next-ai-agents-agentic-era
13. Coyotiv and OpenServ Labs, "Coyotiv and OpenServ Labs Demonstrate Up to 74x AI Reasoning Efficiency Gains in New Research," Newsfile, March 6, 2026. https://www.newsfilecorp.com/release/286412/Coyotiv-and-OpenServ-Labs-Demonstrate-Up-to-74x-AI-Reasoning-Efficiency-Gains-in-New-Research
14. LangChain, "State of Agent Engineering," LangChain, December 2025. https://www.langchain.com/state-of-agent-engineering
