One-line conclusion: Claude Opus 5 scores 61 on the Artificial Analysis Intelligence Index, surpassing GPT-5.6 Sol (59) and Claude Fable 5 (60), but its real value lies in delivering near-frontier intelligence at 26% lower cost per task.
On July 24, 2026, Anthropic officially released Claude Opus 5. Described as "thoughtful and proactive," the model claimed the top spot on the Artificial Analysis Intelligence Index with a score of 61, edging out its sibling Fable 5 (60) and GPT-5.6 Sol (59). More notably, it costs just $2.03 per Intelligence Index task — 26% less than Fable 5 at $2.75.
Opus 5's strength lies in agentic knowledge work. On the GDPval-AA v2 benchmark, Opus 5 (max) scored 1861 Elo — over 100 points ahead of both Fable 5 and GPT-5.6 Sol. On AA-Briefcase, a proprietary agentic knowledge work benchmark, it scored 1720 Elo, +146 ahead of Fable 5. These tests measure the model's ability to produce accurate, professionally-presented outputs using Anthropic's open-source reference agent harness, Stirrup.
However, "most intelligent" doesn't mean "best for every task." The most useful insight in the 2026 AI landscape is task routing — different models excel in different domains.
For frontend design and UI work, Kimi K3 dominates with a 9.5 score on Frontend Code Arena vs Fable 5's 7.5 and Sol's 7.0, at 1/12th the cost of Fable. For backend logic and complex system architecture, Fable 5 leads with SWE-Bench Pro at 80.3%, optimized for long-duration autonomous agent work. Debugging tasks favor GPT-5.6 Sol's iterative, hypothesis-driven approach.
Claude Opus 5's strategy is clear: not replacing Fable 5, but offering "near-frontier intelligence at half the price." Anthropic's pricing implies a two-tier positioning — Fable 5 for the most demanding reasoning-heavy agent tasks, Opus 5 for daily knowledge work and professional document generation.
Notably, NVIDIA CEO Jensen Huang recently published his first public discussion on the importance of open models, arguing that open weights will drive AI ecosystem diversity. Kimi K3 is set to release its full weights on July 27, joining DeepSeek V4 Pro (MIT license, SWE-Bench Verified 80.6%) and GLM-5.2 (1M context) in the rapidly catching open-source camp.
The real question isn't "who's number one," but "which model fits your task." The core skill of 2026 is routing, not brand loyalty.
Frequently Asked Questions (FAQ)
Q: Is Claude Opus 5 really better than GPT-5.6?A: Opus 5 leads the Intelligence Index at 61 vs GPT-5.6 Sol's 59, but GPT-5.6 retains advantages in specific tasks like debugging. There's no "best model," only the best model for each task.
Q: How is Opus 5 different from Fable 5?A: Opus 5 is positioned as "near-frontier intelligence at half the cost" for daily knowledge work; Fable 5 is the top-tier reasoning model for complex autonomous agent tasks, at a higher price.
Q: Which AI model should I use?A: It depends on the task — Kimi K3 for frontend, Fable 5 for complex backend systems, GPT-5.6 Sol for debugging, Opus 5 for daily knowledge work.
Q: What does Claude Opus 5 cost?A: $2.03 per Intelligence Index task on average, 26% below Fable 5 ($2.75) but slightly above Opus 4.8 ($1.80).
Q: Can open-source models catch up to Opus 5?A: Kimi K3 goes open-source on July 27, and DeepSeek V4 Pro already scores 80.6% on SWE-Bench Verified. The open-source camp is closing the gap fast, but top-tier closed models still lead.
#ClaudeOpus5 #AIBenchmarks #GPT56 #KimiK3 #ArtificialIntelligence
留言
張貼留言