RELEASED
2025-11-24
Claude Opus 4.5 breaks 80% on SWE-bench Verified
Anthropic released Claude Opus 4.5, the first model to break 80% on SWE-bench Verified with a score of 80.9%. The launch landed in the middle of a famous two-week stretch in which OpenAI, Google, and Anthropic each shipped flagship coding models days apart.
Opus 4.5 paired the benchmark lead with stronger reasoning and agentic behavior, and Anthropic pitched it as the model for the hardest engineering work. It arrived just as Google's Gemini 3 Pro was resetting expectations elsewhere.
The 80% SWE-bench milestone became a symbolic line in the coding-model race, and Opus 4.5 held it first.