Anthropic shipped Claude Opus 5 this week, and on FrontierBench v0.1 the model scored 43.3% at maximum reasoning effort, ahead of OpenAI's GPT-5.6 Sol at 37.5% on the same benchmark. FrontierBench is one of several evaluation suites tracking frontier-model capability across reasoning-heavy tasks, and the gap reflects performance on that specific test rather than a comprehensive ranking across all workloads.
Benchmark results of this kind measure defined capabilities under controlled conditions, so real-world performance on a given team's actual tasks depends heavily on the specific workload, prompt style, and tooling involved. Independent testing across a broader range of tasks remains the more reliable guide for teams choosing between frontier models, particularly as both companies continue to iterate on their flagship offerings.
The comparison lands amid a broader wave of frontier-model activity this month, with Moonshot AI's Kimi K3 open weights also going live this week and positioning itself as a near-frontier open alternative. Together, the releases underscore how quickly the competitive picture among leading AI labs continues to shift, with benchmark leadership changing hands across successive model releases.