Scores come from the labs themselves or public leaderboards, and benchmarks aren't directly comparable across models. Every benchmark links to its source; lab cards link each score to its source too.
Who leads where
| Benchmark | #1 | #2 | #3 |
|---|---|---|---|
| Terminal-Bench 4.0Agentic coding · Anthropic | #1Claude Opus 5.5 66.4% | #2GPT-6 Astra 57.9% | #3Claude Fable 5.1 55.8% |
| DeepSWE v1.1Software engineering · OpenAI, Tech Insider | #1Muse Spark 1.3 75.4% (max reasoning, via Tech Insider) | #2GPT-6 Astra 74.1% | #3Claude Opus 5 73.7% |
| OSWorld 2.0Computer use · Anthropic, OpenAI (via Kingy AI) | #1Claude Opus 5.5 81.8% (partial credit; different harness than OpenAI's) | #2Claude Fable 5.1 80.7% (partial credit) | #3GPT-6 Astra 72.6% (offline set) |
| SWE-bench VerifiedCoding · vals.ai | #1Claude Opus 5 97.0% | #2GPT-5.6 Sol 96.2% | #3Grok 4.6 95.6% |
| AA Intelligence Index v4.3Overall intelligence · Artificial Analysis | #1Claude Opus 5.5 58 | #2Claude Fable 5.1 53 (tied) | #3GPT-6 Astra 53 (tied) |
| LMArena EloHuman preference · LMArena | #1Claude Mythos 5 1531 (limited access) | #2Claude Fable 5 1525 | #3Claude Opus 5 1522 |
Latest flagship per lab
Also noted
DeepSeek V4 Pro — SWE-bench Verified: 80.6% (Tech Insider)