Artificial Analysis v4.2: Private Test Weight Doubles to 40% to Fight Gaming
Artificial Analysis v4.2 doubles private held-out test weight to 40%, adds two new evals, retires GPQA Diamond. Claude Fable 5.1 leads; GPT-6 Astra tops PDF tasks.

Artificial Analysis published version 4.2 of its Intelligence Index on September 4, doubling the weight of private, held-out test sets from 20% to 40%. The stated reason is direct: the firm wants to reduce labs' ability to optimize models specifically for known benchmarks. Two new evaluations fill that expanded private slot. Claude Fable 5.1 takes the top overall spot on the refreshed leaderboard, with OpenAI's GPT-6 Astra second. The long-running GPQA Diamond benchmark is retired as saturated.
What happened
| Detail | Value |
|---|---|
| Publication date | September 4 |
| Private test-set weight (v4.1) | 20% |
| Private test-set weight (v4.2) | 40% |
| GDP.pdf page count | 4,592 pages across 100 PDFs |
| GDP.pdf domains | 10 |
| Expert-authored grading criteria | 1,275 |
| Retired benchmark | GPQA Diamond (saturated) |
Artificial Analysis describes v4.2 as a stopgap, noting it is “accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier.” The core structural change is that held-out, unpublished test sets now carry twice their previous weight in the overall Intelligence Index score.
The two new evaluations
AA-Briefcase
AA-Briefcase is a private agentic evaluation, meaning models are not just answering questions but working through multi-step tasks. According to Artificial Analysis, it “tests models on realistic agentic knowledge work tasks in complex projects built by industry experts,” with each project spanning many linked tasks and thousands of input source files simulating multi-week work.
GDP.pdf
GDP.pdf was built with Surge AI and focuses on single-turn professional document reasoning. Models must pull and synthesize information spread across 4,592 pages of real PDFs covering text, tables, charts, footnotes, and exclusions. Each answer is graded against 1,275 criteria written by domain experts, across ten professional domains. Because the test set is private, labs cannot tune against it directly.
How the leaderboard changed
The overall ranking puts Claude Fable 5.1 (Anthropic) first, GPT-6 Astra (OpenAI) second, and Meta third. SpaceXAI, Moonshot/Kimi, Z.AI, and Google follow in that order.
The picture shifts on individual slices. GPT-6 Astra posted roughly a 4-point gain over GPT-5.6 Sol to claim second overall, and it leads on the GDP.pdf task specifically. On AA-Briefcase, GPT-6 Astra holds an approximately 85 Elo-point lead over GPT-5.6 Sol.
| Model | GDP.pdf score |
|---|---|
| GPT-6 Astra | 33.2% |
| GPT-5.6 Sol | 28.2% |
| Claude Fable 5.1 | 26.2% |
GPQA Diamond, a graduate-level science benchmark that has featured in AI evaluations for some time, is dropped entirely. Artificial Analysis says it is saturated, meaning top models score so well that it no longer separates them meaningfully.
Why it matters
Benchmark gaming is a real and documented problem in AI development. When labs know exactly which questions are in a test, they can fine-tune toward those specific answers without improving general capability. Raising the private-set share to 40% forces models to generalize, since there is nothing to memorize. It does not eliminate the problem, but it raises the cost of gaming significantly.
The timing adds context. Anthropic’s IPO marketing has reportedly slipped to mid-October. Claude Fable 5.1 sitting at the top of a freshly credible leaderboard is not an irrelevant data point for investors. Similarly, GPT-6 Astra’s strong showing on document and agentic tasks matters for OpenAI’s enterprise positioning. We covered GPT-6 Astra’s availability in GitHub Copilot separately, and document reasoning is exactly the kind of skill that enterprise buyers care about.
For anyone building AI-integrated workflows that touch document-heavy tasks, including legal, finance, or compliance use cases, the GDP.pdf benchmark is worth watching. A model that scores well on 4,592 pages of mixed-format professional documents is closer to production-ready for those jobs than one that scores well on cleaner academic benchmarks.
Our take
The methodology change is the right call, and it is long overdue. Public benchmarks have a shelf life measured in months once frontier labs start optimizing against them. GPQA Diamond being retired is actually a signal of progress: models got too good too fast for it to stay useful. The honest admission that v4.2 is a stopgap before v5 is also refreshing, though it means the leaderboard will shift again soon.
Be cautious about reading the overall rankings as a clean verdict. GPT-6 Astra leads on GDP.pdf and AA-Briefcase, but Claude Fable 5.1 wins overall. The gap depends heavily on which tasks you weight. If you are choosing a model for agentic document work specifically, the per-task numbers matter more than the composite score. We see this play out with clients regularly: the best general-purpose model is rarely the best model for a specific workflow.
If you are not yet tracking how different frontier models perform on your actual tasks rather than public leaderboards, that is the practical gap to close. Our recent piece on AI workflow pitfalls covers why benchmark scores often mislead in real production settings.
What to do about it
- Check whether any internal model-selection decisions were based on GPQA Diamond scores. That benchmark is now retired and no longer a useful reference point.
- If your use case involves professional document analysis (contracts, financial reports, compliance filings), compare GDP.pdf scores across the models you are evaluating. GPT-6 Astra’s 33.2% versus Claude Fable 5.1’s 26.2% is a meaningful spread for that workload.
- For agentic, multi-step knowledge work, weight the AA-Briefcase results. The roughly 85 Elo-point gap between GPT-6 Astra and GPT-5.6 Sol on that task is worth testing against your own workflows.
- Wait for v5 before locking in any long-term model strategy. Artificial Analysis has already flagged that a fuller release is coming.
The practical takeaway: pick models by running them on a sample of your real work, not on a composite leaderboard score that blends tasks you will never use.
Frequently asked questions
What changed in Artificial Analysis Intelligence Index v4.2?
V4.2, published September 4, doubles the weight of private held-out test sets from 20% to 40%, adds two new evaluations (AA-Briefcase and GDP.pdf), and retires the GPQA Diamond benchmark as saturated.
What is the GDP.pdf benchmark?
GDP.pdf is a new evaluation built with Surge AI that tests professional document reasoning across 100 PDFs spanning 4,592 pages and ten domains. Answers are graded against 1,275 expert-authored criteria.
Which AI model leads the Artificial Analysis v4.2 leaderboard?
Claude Fable 5.1 by Anthropic ranks first overall. GPT-6 Astra by OpenAI is second and leads on the GDP.pdf and AA-Briefcase individual tasks.
Why is GPQA Diamond being retired from AI benchmarks?
Artificial Analysis retired GPQA Diamond because it is saturated. Top frontier models now score so highly on it that the benchmark can no longer meaningfully distinguish between them.


