Anthropic's Claude Opus 5 and OpenAI's GPT-5.6 Sol represent the two most capable coding agents available in mid-2026. Opus 5 launched on July 24, 2026, two weeks after GPT-5.6 Sol's July 9 release. Both target the same high-stakes use case: autonomous software engineering that produces maintainer-grade changes rather than merely passing tests. The aggregate leaderboard tells a clear story at the top: Claude Opus 5 (Adaptive Reasoning, Max Effort) ranks #1 with a score of 61, followed by Opus 5 (Xhigh Effort) at 60, Claude Fable 5 at 60, and GPT-5.6 Sol (max) at 59.
Pricing is the first practical difference. Opus 5 charges $5.00 per million input tokens and $25.00 per million output tokens. GPT-5.6 Sol charges $5.00 per million input tokens and $30.00 per million output tokens, making Opus 5 roughly 17% cheaper on output. Cached input costs $0.50 per million tokens for both. Both models offer a 50% batch API discount. Opus 5 also provides a fast mode that runs at 2.5× speed for 2× price, while GPT-5.6 Sol has no dedicated fast mode. Context windows are comparable: Opus 5 supports 1,000,000 tokens, Sol supports 1,050,000 tokens. Both cap maximum output at 128,000 tokens. Knowledge cutoffs differ slightly: Opus 5's training knowledge extends to January 2026, while Sol's extends to February 16, 2026.
On shared benchmarks, Opus 5 leads 9–3. The most dramatic gap is SWE-bench Pro, where Opus 5 scores 79.2% against Sol's 64.6%—a 14.6-point margin. SWE-bench Verified is closer: Opus 5 at 96.0% versus Sol at 95.0%. On Frontier-Bench v0.1, which measures agentic coding on novel problems, Opus 5 reaches 43.3% while Sol reaches 34.4%. Cognition's FrontierCode 1.1, which evaluates whether a change would actually be merged by maintainers rather than simply passing tests, gives Opus 5 53.4% mergeability versus Sol's 47.5%, with pass rates of 58.9% and 52.9% respectively. FrontierCode also reports lower observed cost per rollout for Opus 5: $4.30 versus Sol's $6.30, despite nearly identical output token counts (33.6K versus 33.2K). Notably, Opus 5 achieved this FrontierCode result at medium effort, while Sol ran at max effort.
The following table summarizes the head-to-head numbers that matter most for agentic coding:
| Metric | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|
| Leaderboard rank | #1 (score 61, Max Effort) | #4 (score 59, max) |
| Input price | $5.00 / 1M tokens | $5.00 / 1M tokens |
| Output price | $25.00 / 1M tokens | $30.00 / 1M tokens |
| Cached input | $0.50 / 1M tokens | $0.50 / 1M tokens |
| Context window | 1,000,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| SWE-bench Pro | 79.2% | 64.6% |
| SWE-bench Verified | 96.0% | 95.0% |
| Frontier-Bench v0.1 | 43.3% | 34.4% |
| FrontierCode 1.1 mergeability | 53.4% | 47.5% |
| FrontierCode 1.1 pass rate | 58.9% | 52.9% |
| Terminal-Bench 2.1 | not reported | 91.9% (with Ultra sub-agents) |
| DeepSWE v1.1 | not reported | 72.7% |
| BrowseComp | 90.8% | 90.4% |
| ARC-AGI-3 | 3.9× next best | not reported |
| GDPval-AA v2 Elo | 1861 | 1736 |
| AA-Briefcase Elo | 1720 | 1505 |
| HealthBench Professional | 59.8% | 60.5% |
| AutomationBench | 26.0% | 18.1% |
| BenchAlign Score | 85.88 (#1 overall) | 81.46 (#4 overall) |
GPT-5.6 Sol wins where the task rewards terminal fluency and deep repository navigation. Its Terminal-Bench 2.1 score of 91.9% with Ultra sub-agents is the standout number for DevOps, system administration, and command-line automation. Sol also leads on DeepSWE v1.1 at 72.7% and edges out Opus 5 on HealthBench Professional, 60.5% to 59.8%. On BrowseComp the two are effectively tied: Opus 5 at 90.8% and Sol at 90.4%. Anthropic's own system card, which cites OpenAI's published figures for GPT-5.6 Sol, additionally notes that Sol leads on ARC-AGI-2, while Opus 5 leads on ARC-AGI-3 with a 3.9× advantage over the next best result.
Speed is Sol's other advantage. Through standard Anthropic API deployment, Opus 5 reaches approximately 59.8 tokens per second at max effort, while Sol reaches approximately 53.8 tokens per second through standard OpenAI deployment. However, Sol on Cerebras hardware pushes 750 tokens per second, making it the clear choice when raw throughput dominates the decision.
Availability may be the decisive factor for many teams. Claude Opus 5 is generally available through the Anthropic API, Claude Pro, Claude Max, OpenRouter, and Claude Code. GPT-5.6 Sol is government-gated, restricted to approved organizations, and in some cases requires security clearance. For most commercial developers, Sol is not currently an option regardless of its benchmark strengths. OpenAI's generally available alternative in the same family is GPT-5.6 Terra at $3.00 per million input tokens and $9.00 per million output tokens, though Terra competes more directly with Claude Sonnet 5 than with Opus 5.
Reasoning modes differ as well. Opus 5 uses adaptive reasoning that is always on, with effort ranging from low to max. Sol uses a Pro mode with max and ultra settings, where ultra can deploy sub-agents. Both models accept text and image input and produce text output.
For agentic coding specifically, the evidence points toward Opus 5 as the safer default. It wins on the benchmarks most closely tied to real software engineering—SWE-bench Pro, SWE-bench Verified, Frontier-Bench, and FrontierCode mergeability—while costing less per output token and per rollout. Its BenchAlign Score of 85.88 ranks #1 overall, compared with Sol's 81.46 at #4. Sol remains the better pick for terminal-heavy infrastructure work and for deployments where Cerebras-level speed is available and access restrictions can be satisfied. If neither access nor specialized terminal work is the deciding constraint, Opus 5 currently offers the stronger combination of mergeable output, breadth, and price.