ChengRang

Artificial Analysis

AI Search & Research Free

Independent third-party AI model benchmarking site; the Intelligence Index weights nine evaluations, and v4.2 raised the private held-out share to 40% to curb leaderboard gaming

Model BenchmarkingLeaderboardIndependentCost Comparison
Visit Artificial Analysis

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Artificial Analysis is an independent model benchmarking site, best known for the Intelligence Index, a composite score built from weighted evaluations. Version 4.2 landed on September 4, 2026 as an interim release ahead of v5, pulling forward parts of the v5 plan.

It also publishes cost-per-task and output token efficiency curves alongside the score, which is why many teams use it as the starting point for model selection.

Key Features

Use Cases

Pros

Pricing

The site is free to browse. Methodology and per-model breakdowns are published at artificialanalysis.ai/methodology/intelligence-benchmarking.

Summary

A leaderboard is only worth as much as it is hard to feed. Nearly every v4.2 change points the same direction — shrinking the room labs have to optimise against known questions. Doubling the private share, retiring the saturated GPQA Diamond and patching the grading sandbox all push scores closer to real performance on unfamiliar work. In v4.2, Claude Fable 5.1 leads overall, with GPT-6 Astra second and first on GDP.pdf at 33.2%. Note that scores shift with each index version, so comparing raw numbers across versions is misleading.

Version History

Category
AI Search & Research
Pricing
Free
Tags
Model Benchmarking · Leaderboard · Independent

Related Tools