On July 24 (local time), Anthropic released a new model, Claude Opus 5. The moment it launched, it topped the Intelligence Index run by the independent evaluator Artificial Analysis, scoring about 61 points. What stands out is the price. It holds the same $5 input / $25 output per million tokens as the prior generation, Opus 4.8, yet its SWE-bench score โ€” a measure of software development ability โ€” leapt a full generation. Because the announcement landed in the middle of a July โ€œmodel showdownโ€ in which major AI companies rolled out new models one after another, today we lay out that broader flow. ๐Ÿ”

TL;DR: Claude Opus 5, released July 24, took first place on both the Artificial Analysis Intelligence Index (about 61) and the Agentic Index (55.3). Its price is unchanged at $5 input / $25 output per million tokens, yet coding jumped sharply, with 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro. In July, OpenAI (the GPT-5.6 family) and Google (Gemini) also shipped new models, kicking a โ€œprice-per-intelligenceโ€ competition into high gear.

๐Ÿ† What Happened โ€” the Intelligence Crown Changed Hands Again

The core of the news is that Claude Opus 5 reached the top of an independent benchmark. Artificial Analysis is an outfit that bundles many test items to score a modelโ€™s overall intelligence. Opus 5 scored about 61 on its Intelligence Index (some tallies list it as 60.7%) to take the lead. It also ranked first on the Agentic Index โ€” which measures the ability to pick tools on its own and handle multi-step tasks โ€” with 55.3.

On both metrics, it edged out the same companyโ€™s higher model, Fable 5, as well as OpenAIโ€™s GPT-5.6 Sol. That said, rankings flip whenever a new model appears, so it is better to read this as a flow โ€” top-tier competition swinging on a scale of weeks โ€” rather than fixating on โ€œhow many points, what rank.โ€

๐Ÿ’ฐ Same Price, a Full Generation of Performance โ€” โ€œPrice-per-Intelligenceโ€ Is the Point

The truly eye-catching part of this release is the price. Opus 5 is priced at $5 input / $25 output per million tokens, exactly the same as the previous model, Opus 4.8. Normally, when performance rises, price rises with it โ€” but this time performance was lifted while the price stayed put.

Anthropic says Opus 5 delivers intelligence close to its higher-tier Fable 5 at half the cost (Fable 5 runs $10 input / $50 output). It also introduced a โ€œfast modeโ€ roughly 2.5ร— quicker than the default, priced at $10 input / $50 output per million tokens. It is a signal that the center of gravity in competition is shifting toward the same performance for less, or better performance at the same price.

๐Ÿ‘จโ€๐Ÿ’ป The Threshold Agentic Coding Just Crossed โ€” SWE-bench 96%

For developers, the coding numbers are what to watch most. On SWE-bench Verified โ€” which has models fix real open-source bugs โ€” Opus 5 scored 96.0%, and on the tougher SWE-bench Pro, 79.2%. SWE-bench Pro in particular climbed about 10 percentage points, from 69.2% on the prior-generation Opus 4.8 to 79.2%.

On Frontier-Bench v0.1, which gauges โ€œagentic codingโ€ โ€” working through multiple steps on its own โ€” Opus 5 scored 43.3%, roughly double Opus 4.8โ€™s 21.1%. That means the capability is broadening beyond writing a single good line of code toward carrying a task to completion across many files with less human intervention. Opus 5 became available immediately on the Claude API as well as major clouds like Amazon Bedrock, Google Cloud, and Microsoft Foundry, and in the Claude app and Claude Code.

๐ŸŒŠ Julyโ€™s โ€œAI Model Showdownโ€ โ€” How the Competitive Terrain Is Flowing

Opus 5 is also one scene in a new-model race that ran all month. According to various sources, OpenAI formally released the GPT-5.6 family (the Sol, Terra, and Luna tiers) early in July, and Google is reported to have introduced a lightweight Gemini model in late July. Google, however, is said to have somewhat delayed a wide release of its top-tier frontier model based on internal evaluation results โ€” a point best treated as observation until each company confirms it officially.

The common thread is a clearer trend of each company shipping models split into detailed tiers (a fast model, a low-cost model, a top-tier model). Alongside contesting peak performance, competition is also spreading toward offering the same performance at a lower price.

๐Ÿ‡ฐ๐Ÿ‡ท What It Means for Companies and the Field โ€” Faster โ€œAI Agentโ€ Adoption

The message this flow sends to real work is that โ€œAI agentsโ€ are spreading fast. Agent-style AI โ€” which judges on its own and handles multi-step tasks โ€” is settling into coding, documents, and workflow automation. Research firm Gartner has projected that by the end of 2026, about 40% of enterprise applications will carry AI agent features.

If the pattern of rising performance with flat or falling prices continues, companies can run more AI work on the same budget. That lowers the bar for AI adoption at domestic firms and dovetails with the large-scale AI data center and semiconductor investment demand announced earlier. Still, a high benchmark score does not immediately guarantee real business results, so it is wise to run a verification pass fitted to your own work when adopting.

๐Ÿ“Œ The Takeaway โ€” Watch the Flow, Not the Numbers

This weekโ€™s AI model news comes down to three points. First, frontier competition is hot enough that the top intelligence ranking flips on a scale of weeks. Second, the axis of competition is shifting from โ€œthe smartest modelโ€ toward price-per-performance โ€” โ€œsmarter at the same price, or the same performance for less.โ€ Third, as coding and agent performance leaps a generation at a time, there is more room for real-world work automation to accelerate.

That said, benchmark scores and rankings are relative measures that change with every new model, and each companyโ€™s measurement standards differ slightly. Rather than hastily declaring one model superior, this is a moment that calls for weighing performance, price, stability, security, and fit with your own data all together.

โ€ป This article is for informational purposes only and is not investment advice.

Sources