Overview
Artificial Analysis is an independent model benchmarking site, best known for the Intelligence Index, a composite score built from weighted evaluations. Version 4.2 landed on September 4, 2026 as an interim release ahead of v5, pulling forward parts of the v5 plan.
It also publishes cost-per-task and output token efficiency curves alongside the score, which is why many teams use it as the starting point for model selection.
Key Features
- Nine evaluations, one composite: The Index weights four categories: Agents at 34%, Coding at 24%, Scientific Reasoning at 24% and General at 18%.
- Private test share doubled to 40%: Up from 20% in v4.1, drawing on AA-Briefcase, AA-Omniscience and CritPt solutions. Public benchmarks tend to leak into training data, so raising the private share directly increases the cost of gaming the leaderboard.
- AA-Briefcase added: An in-house agentic knowledge-work evaluation where industry experts design projects spanning weeks, with interlinked tasks and thousands of source files. Models run in a sandboxed Linux environment with no internet access, up to 500 turns and a code execution tool. Grading combines binary rubric checks with pairwise Elo comparisons on analytical and presentation quality, judged by a rotating panel of Opus 4.8, GPT-5.5 and Gemini 3.1 Pro Preview to reduce same-family bias.
- GDP.pdf added: Built by Surge AI, this tests single-turn professional document reasoning across 100 PDFs and ten domains, spanning 4,592 pages of prose, tables, charts, footnotes and exclusions. Answers are graded against 1,275 expert-authored atomic criteria, and the headline All-pass Rate credits a task only when every criterion is satisfied.
- GPQA Diamond retired: The graduate-level science multiple-choice test has been saturated by frontier reasoning models and no longer separates them, so it was dropped entirely.
- Grading infrastructure hardened: AA-LCR moved to v1.1 with a grading system prompt and corrections to errors and ambiguities in answer keys; GDPval-AA v2 and AA-Briefcase improved sampling and re-anchored the Elo scale; SciCode relaxed sandbox execution limits so slow but correct code no longer counts as a failure.
- Cost shown alongside score: Beyond the score, cost-per-task and output token efficiency curves are published together, so selection can be driven by budget rather than rank alone.
Use Cases
- Comparing frontier models on score, cost and token efficiency before committing
- Tracking one lab's release cadence and relative position over time
- Judging real completion rates on agentic work rather than single-question scores
- Citing third-party benchmark data with a public, traceable source
Pros
- A high private test share makes scores harder to tune against known questions
- Cost and token efficiency sit next to the score, so trade-offs are visible
- Methodology and per-model breakdowns are published
- Retiring saturated benchmarks keeps the leaderboard discriminating
Pricing
The site is free to browse. Methodology and per-model breakdowns are published at artificialanalysis.ai/methodology/intelligence-benchmarking.
Summary
A leaderboard is only worth as much as it is hard to feed. Nearly every v4.2 change points the same direction — shrinking the room labs have to optimise against known questions. Doubling the private share, retiring the saturated GPQA Diamond and patching the grading sandbox all push scores closer to real performance on unfamiliar work. In v4.2, Claude Fable 5.1 leads overall, with GPT-6 Astra second and first on GDP.pdf at 33.2%. Note that scores shift with each index version, so comparing raw numbers across versions is misleading.
Version History
- Intelligence Index v4.2 (2026-09-04): Private test share raised from 20% to 40%; AA-Briefcase and GDP.pdf added; GPQA Diamond retired
- Intelligence Index v4 (2026-01): Composite scoring rebuilt and the four-category weighting fixed