Skip to content

Claude vs Codex vs DeepSeek vs Gemini: Code, Autotests, Business Plans & Reasoning (2026)

Updated 2026-07-26

Updated 26 July 2026, two days after Claude Opus 5 shipped. Opus 5 now tops the independent Artificial Analysis Intelligence Index (61 at max effort) and the SWE-bench Verified leaderboard (97.0%), narrowly ahead of OpenAI's GPT-5.6 Sol (59 index, 96.2% SWE-bench) — a much tighter race than the raw ranking suggests. Gemini's flagship 3.5 Pro has now slipped three times and is still not generally available, so 3.1 Pro remains Google's shipping model. DeepSeek V4-Pro is still 5-10x cheaper than anything else here and the only self-hostable option, but it now sits far below the frontier on the re-weighted, agentic-heavy index. Every vendor's own numbers are higher than what independent evaluators measure — treat that gap as the headline finding, not a footnote.

Most "Claude vs GPT vs Gemini" pages on the internet just paste a vendor's own launch-blog table and call it a comparison. This one is built differently: vendor claims are cross-checked against independent trackers — Artificial Analysis, NIST's CAISI, Epoch AI, vals.ai's SWE-bench leaderboard, LMArena — specifically for four use cases QA and dev teams actually care about: writing code, writing and running automated tests, drafting business plans and analytical documents, and working across large, long-lived codebases. This July 2026 refresh covers Anthropic's Claude (Opus 5, released 24 July 2026), OpenAI's Codex (running on GPT-5.6 Sol, generally available since 9 July 2026), DeepSeek (V4-Pro), and Google's Gemini (3.1 Pro — still the shipping flagship, because Gemini 3.5 Pro has missed three announced dates). Every number below is labeled vendor-reported or independent, because that distinction changed the ranking more than any single benchmark did.

Key takeaways

  • As of 26 July 2026, Claude Opus 5 leads both the independent Artificial Analysis Intelligence Index (61 at max effort) and vals.ai's SWE-bench Verified leaderboard (97.0%) — but GPT-5.6 Sol is within 0.8 points on SWE-bench, so treat them as tied on coding and decide on cost, latency, and ecosystem instead.
  • Opus 5 replaced Opus 4.8 at identical pricing ($5 / $25 per MTok) with a 1M-token window that is both default and maximum, billed flat — while GPT-5.6 Sol doubles to $10 / $45 per MTok past 272K tokens, which agentic runs hit routinely.
  • Google's Gemini 3.5 Pro has now missed three announced dates and was reportedly delayed over hallucination rates and reliability; 3.1 Pro remains the shipping flagship and only Flash-tier 3.5/3.6 models launched this quarter.
  • Artificial Analysis re-weighted its Intelligence Index to v4.1 around agentic workloads in 2026 — index scores quoted earlier in the year are measured on a different yardstick and are not comparable to the numbers on this page.
  • DeepSeek V4-Pro is roughly 30x cheaper per output token than Opus 5 and the only self-hostable option, but scores 44 on the re-weighted index and confidently fabricates answers 94-96% of the time when it doesn't know (AA-Omniscience) — a specific risk for analytical writing.
  • Several of Opus 5's launch claims are published only as ratios against unnamed competitors ("3× the next-best" on ARC-AGI 3, "~1.5×" on AutomationBench) with no absolute score — ratios without denominators can't be independently checked and aren't treated as numbers here.

At a glance

Claude (Opus 5)Codex (GPT-5.6 Sol)DeepSeek (V4-Pro)Gemini (3.1 Pro)
Released / current as of24 Jul 20269 Jul 2026 (GA)Apr 2026Shipping flagship; 3.5 Pro delayed again 16 Jul 2026
Price (input / output per MTok)$5 / $25$5 / $30 (≤272K ctx); $10 / $45 beyond$0.44 / $0.87$2 / $12 (≤200K ctx)
Context window1M tokens (default and maximum), 128K max output1.05M tokens, 128K max output; 400K inside the Codex product1M tokens1M tokens
Open weights / self-hostableNoNoYes (MIT license)No
Code — SWE-bench Verified (vals.ai independent leaderboard)97.0% (#1)96.2%80.6% (vendor-reported)80.6% (vendor-reported)
Code — Terminal-Bench 2.x (agentic CLI tasks)Component of the AA Intelligence Index v4.1; no standalone figure published at launch82.7% (v2.0, one tracker)67.9% (v2.0)76.2% (v2.1, vendor)
Code — real-world developer sentimentPraised for finishing multi-file work in one pass; AA measures it as notably slow and verboseStrong on terminal/DevOps, weaker on frontend/UI (community)"Best value, not best coder" — good for volume, not hardest bugs (dev reviews)Benchmark-strong but described as over-engineered/verbose in real use (HN)
Autotests — independent hands-on review23/23 passing tests, 95% line coverage on the Opus 4.8-era review, but missed real edge cases (ontestautomation.com)Can run/iterate tests autonomously in cloud sandbox; reports of over-mocking / happy-path biasCommunity tutorials only; Flash tier reportedly lags Pro on multi-step debug loopsFirst-party IDE test generation (Android Studio); HN flags agentic harness-escape issues
Autotests — agentic tool-use benchmark proxyOSWorld 2.0: beats Fable 5 at ~⅓ the cost (vendor); Zapier AutomationBench pass rate ~1.5× next-best (vendor)OSWorld-Verified: SOTA claimed (vendor)Terminal-Bench 2.0: 67.9%MCP Atlas: 83.6% (vendor)
Business writing — GPQA Diamond (PhD-level reasoning proxy)Component of AA Intelligence Index v4.1; no standalone figure published at launch93.6%90.1% (vendor)94.3% (independent-adjacent, DeepMind card)
Business writing — documented fabrication/hallucination riskLowest misaligned-behaviour score of any recent Claude (2.3, vendor safety eval); no independent hallucination figure yetNo specific figure found in this pass94-96% hallucination rate ON questions it gets wrong (AA-Omniscience) — confidently wrong rather than abstaining3.5 Pro was reportedly delayed specifically over hallucination rates (Bloomberg, 16 Jul 2026)
Business writing — native office/workspace integrationNo native suite; Claude Cowork + long-context APINo native suite; ChatGPT app + APINo native suiteNative in Gmail/Docs/Sheets/Slides (Workspace)
Reasoning — mechanismAdaptive thinking on by default, effort dial low→max; raw chain of thought never returnedSelectable reasoning effort low→xhighHybrid thinking/non-thinking per requestDeep Think = extra inference-time compute, separate mode
Reasoning — competition math (AIME/FrontierMath)No standalone figure published at launchFrontierMath Tier 4: ~31-40% (Epoch AI, independent, GPT-5.5-era)OTIS-AIME-2025: 97% (NIST CAISI, independent)Reported behind GPT-5.4 on AIME; no confirmed figure
Context — long-context retrieval at full window (MRCR-style, degradation check)1M is default and maximum, flat pricing, no long-context surcharge; vendor multi-needle claims unverifiedMRCR 1M: 74.0% (GPT-5.5-era); long-context requests above 272K are billed at double rateMRCR 8-needle: 0.82 @256K → 0.59 @1M (real drop-off, vendor)MRCR: 84.9% @128K → 26.3% @1M (steep drop-off, vendor model card)
Independent composite — Artificial Analysis Intelligence Index v4.161 (max effort) — #1 of ~190 models59 (max)44 (max)Not separately published; Gemini 3.5 Flash (high) scores 50
Independent — measured output speed52.6 tok/s, ~68s to first token (AA) — among the slowest frontier modelsNot captured in this passNot captured in this passNot captured in this pass
Known risk / honest caveatOpus 5 trails Anthropic's own Mythos 5 on offensive-cybersecurity tasks and ships tighter cyber guardrails — benign security work can trip a refusalLong-context pricing doubles past 272K tokens, which is easy to hit in agentic runsHosted app stores data in mainland China (governed by Chinese law); banned on gov't systems in several countries3.5 Pro has slipped three times and the base model was reportedly scrapped and rebuilt; only Flash-tier 3.5/3.6 models have shipped

Independent intelligence ranking (July 2026)

Artificial Analysis re-weighted its Intelligence Index to v4.1 in 2026, shifting it toward agentic workloads (GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR). That rescaling means these numbers are not comparable to index scores quoted earlier in the year — models did not all drop, the yardstick changed. Scores are configuration-specific: the same model scores differently at different reasoning-effort settings, which is why the effort level is part of each label.

Independent intelligence ranking (July 2026)Intelligence Index v4.1 (0-100)
ModelIntelligence Index v4.1 (0-100)
Claude Opus 5max effort61
Claude Opus 5xhigh effort60
Claude Fable 5max effort, Opus 4.8 fallback60
GPT-5.6 Solmax59
Claude Opus 5high effort59
GLM-5.2max — highest-ranked open weights51
Gemini 3.5 Flashhigh50
DeepSeek V4-Promax effort44
Median of all tracked models32

Source: Artificial Analysis Intelligence Index, snapshot 26 July 2026

What a million output tokens costs

Intelligence ranking and cost ranking are close to inverted, which is the whole trade-off. Note two asymmetries the sticker price hides: GPT-5.6 Sol jumps to $45 per million output tokens on requests past 272K context, and Claude's 1M window is billed flat with no long-context surcharge. DeepSeek's bar is not a rounding error — at $0.87 it is roughly 30x cheaper than Opus 5 per output token.

What a million output tokens costsUSD per 1M output tokens (list price)
ModelUSD per 1M output tokens (list price)
Claude Fable 550
GPT-5.6 Sollong context, >272K45
GPT-5.6 Solstandard context30
Claude Opus 525
Gemini 3.1 Pro≤200K context12
Claude Sonnet 5intro pricing through 31 Aug 202610
DeepSeek V4-Pro0.87

Source: Vendor pricing pages, checked 26 July 2026

What changed in July 2026

This comparison moved substantially in a single month, which is the best argument for treating any fixed table as a snapshot. Anthropic shipped Claude Opus 5 on 24 July 2026 at the same $5 / $25 per MTok as the Opus 4.8 it replaces — an unusual launch in that the price did not move while the independent index score did, taking the top slot at 61. OpenAI made the GPT-5.6 family (Sol, Terra, Luna) generally available on 9 July, with Sol as the flagship behind Codex and the `gpt-5.6` API alias. Google, meanwhile, did not ship its flagship at all: Gemini 3.5 Pro missed a June target, then a widely-reported 17 July date, with Bloomberg reporting on 16 July that the release slipped again after the model fell short of Google's internal quality bar on hallucination rates and real-world reliability — DeepMind reportedly scrapped and rebuilt the base model. Google did ship Gemini 3.6 Flash, 3.5 Flash-Lite, and a security-focused 3.5 Flash Cyber. Separately, Artificial Analysis re-weighted its Intelligence Index to v4.1 around agentic workloads, so index numbers quoted anywhere before that change are measured on a different yardstick and should not be compared to the ones on this page.

Why vendor benchmarks and independent benchmarks disagree

Every lab publishes SWE-bench Verified, GPQA Diamond, or Terminal-Bench numbers under its own best-case configuration: highest reasoning effort, full tool access, and often multiple attempts averaged together. Independent evaluators — Artificial Analysis, NIST's CAISI, Epoch AI's FrontierMath, vals.ai's SWE-bench leaderboard — reproduce the same benchmark under one fixed harness, and the numbers are consistently lower. The gap isn't noise: NIST CAISI measured DeepSeek V4 at 74% on comparable software-engineering tasks against DeepSeek's own ~80.6% SWE-bench claim, and separately found its real-world capability trailing the frontier by roughly eight months. Anthropic's Claude Fable 5 had its SWE-bench Pro claim publicly disputed by evaluators within 24 hours of its June launch. The Opus 5 launch is a useful contrast: several of its headline claims are stated only as ratios against unnamed competitors ("three times as high as the next-best model" on ARC-AGI 3, "around 1.5×" on Zapier AutomationBench, "more than doubles Opus 4.8" on Frontier-Bench v0.1) with no absolute score published. Those may well be true, but a ratio without a denominator cannot be independently checked, which is why none of them appear as a number in the table above.

Code writing: how they actually compare

On vals.ai's independently-run SWE-bench Verified leaderboard, Claude Opus 5 currently leads at 97.0%, with GPT-5.6 Sol at 96.2% and Claude Fable 5 at 95.0%. An 0.8-point gap on a benchmark this saturated is not a decisive lead — at the top of SWE-bench the remaining failures are increasingly ambiguous or badly-specified tasks, so treat Opus 5 and GPT-5.6 Sol as effectively tied on this measure and pick on the other axes. Neither DeepSeek V4-Pro nor Gemini 3.1 Pro has a comparable independent SWE-bench figure at this tier; their ~80.6% vendor-reported numbers are a generation behind. Where the two frontier models genuinely differ is cost shape and speed: Codex's context pricing doubles past 272K tokens, which agentic runs hit routinely, while Claude prices its full 1M window flat. In the other direction, Artificial Analysis measures Opus 5 at 52.6 output tokens per second with roughly 68 seconds to first token, and explicitly characterises it as slow and verbose — for an overnight refactor that is irrelevant, for an interactive inner loop it is the thing you will notice first. Developer sentiment continues to describe Claude's agentic style as the one most likely to finish multi-file work in a single pass, Codex as strongest on terminal and DevOps workflows but weaker on frontend, Gemini's output as powerful but over-engineered, and DeepSeek as good value rather than best-in-class on the hardest bugs.

Autotests and test automation: how they actually compare

No vendor publishes a dedicated "writes good tests" benchmark — this remains qualitative territory for all four models. The most useful independent data point is still a hands-on QA-practitioner review (not vendor-affiliated) of Claude Code from the Opus 4.8 era: it generated a 23-test Java suite with 95% line coverage and 91% mutation coverage in minutes, but missed real edge cases (an HTTP 500 path, a boundary condition on interest calculations) and produced some redundant tests. That review has not been re-run on Opus 5, so treat it as evidence about the workflow rather than about the current model. The nearest proxies for Opus 5 specifically are agentic-environment benchmarks, where Anthropic reports it beating Fable 5 on OSWorld 2.0 at about a third of the cost and posting roughly 1.5× the next-best pass rate on Zapier's AutomationBench — both vendor-reported. Codex's advantage remains architectural: it runs tests inside a sandboxed cloud container and iterates autonomously before opening a PR, though community reports flag over-mocking and happy-path bias. Gemini ships first-party test generation inside Android Studio and Code Assist, but Hacker News threads on its agentic CLI environments flag harness-escape and reliability issues relevant to unattended CI. DeepSeek's test-automation evidence remains community tutorials rather than evaluation. Across all four, the honest guidance has not changed: model output is a fast first draft that still needs a human test-strategy review, and coverage percentage is the least informative number in the report.

Business plans and analytical writing: how they actually compare

For long-form business documents, the most actionable finding still isn't a benchmark score — it's a documented failure mode. DeepSeek V4-Pro scores respectably on GPQA Diamond (90.1%, vendor), but Artificial Analysis's AA-Omniscience benchmark found that when it doesn't know an answer it guesses confidently wrong 94-96% of the time rather than admitting uncertainty. That is a real risk for stakes-bearing business writing, and it is the single most useful thing to know before deploying the cheap option on analytical work. Hallucination is also, notably, the reason Google's flagship is late: the reported basis for Gemini 3.5 Pro's July delay was falling short on hallucination rates and real-world reliability, not on capability benchmarks. Anthropic reports Opus 5 with the lowest misaligned-behaviour score of any recent Claude (2.3 on its own safety evaluation) and materially better scientific-domain accuracy than Opus 4.8 — 10.2 percentage points higher on organic chemistry and 7.7 on protein-related tasks — though both are vendor-measured and neither is a general hallucination rate. Gemini's differentiator here remains structural rather than numeric: native integration into Gmail, Docs, Sheets, and Slides, letting it draft directly from existing Workspace files, which none of the other three can do natively.

Analytical reasoning approach: how they actually compare

The four labs still solve reasoning differently, and the mechanism now matters as much as the score. Claude's Opus 5 runs adaptive thinking by default — a change from Opus 4.8, where omitting the thinking parameter meant no thinking — with a five-step effort dial from low to max; the effort setting alone moves its index score from 59 to 61, which is a wider spread than the gap to the next vendor. Its raw chain of thought is never returned, only summaries, so reasoning-transparency workflows have to be rebuilt around that. OpenAI exposes a comparable low-to-xhigh reasoning-effort control across the GPT-5.6 tiers. Gemini keeps "Deep Think" as a separate mode that spends extra inference-time compute rather than switching to a bigger model, and still posts the highest GPQA Diamond figure in this group (94.3%). DeepSeek's RL-driven reasoning lineage shows real, independently-verified strength on competition math (97% on OTIS-AIME-2025 per NIST CAISI) alongside an independently-measured weakness on novel abstract reasoning. The practical takeaway for anyone benchmarking these models themselves: report the effort setting alongside the score, because a comparison that doesn't state it isn't reproducible.

Context retention and project understanding: how they actually compare

All four advertise roughly 1M-token context windows, but "1M tokens" and "reliable recall at 1M tokens" remain different claims. Gemini 3.1 Pro's own model card shows MRCR retrieval accuracy collapsing from 84.9% at 128K tokens to 26.3% at the full 1M mark. DeepSeek shows the same pattern at smaller scale (0.82 at 256K down to 0.59 at 1M) but backs it with real architectural efficiency gains — a sparse-attention design cutting KV-cache memory to roughly a tenth of its prior generation at 1M scale. Codex has two separate gotchas: GPT-5.6 Sol supports 1.05M tokens via the API but only 400K inside the Codex product, and API requests above 272K tokens are billed at $10 / $45 per MTok instead of $5 / $30 — so the long-context path is both narrower and more expensive than the headline number implies. Claude's Opus 5, Sonnet 5, and Fable 5 all ship 1M tokens at flat per-token pricing with no long-context surcharge, which remains a genuine and verifiable differentiator; on Opus 5, 1M is both the default and the maximum, so there is no smaller variant to accidentally land on via the API. The vendor's specific multi-needle retrieval claims still could not be independently confirmed.

Which model should you actually pick

For day-to-day coding and test automation, Claude Opus 5 is the strongest single pick in this comparison — top of the independent intelligence index, top of the independent SWE-bench leaderboard, flat 1M-token pricing, and no price increase over the model it replaces. The caveats are latency (measurably slow to first token, and verbose) and tighter cybersecurity guardrails that can refuse benign security-adjacent work. If your team is standardised on GitHub and wants autonomous PR-open-and-review workflows, Codex on GPT-5.6 Sol is within a point of Claude on SWE-bench and is arguably the deeper agentic product — budget for the 400K in-product context ceiling and the doubled long-context rate past 272K. If cost is the binding constraint, or you need self-hosting for compliance, DeepSeek V4-Pro is roughly 30x cheaper per output token than Opus 5 and the only genuinely open-weight option, at the price of sitting well below the frontier on the agentic-weighted index and having a documented tendency to fabricate rather than abstain. If your organisation runs on Google Workspace, Gemini 3.1 Pro remains the natural fit — but plan around the fact that its successor has now missed three dates and only Flash-tier models shipped this quarter. Within Claude's own lineup, Sonnet 5 keeps most of Opus 5's coding ability at roughly a third of the price for teams that don't need the ceiling, and Fable 5 at $10 / $50 is now hard to justify for coding specifically, since Opus 5 matches or beats it on several of Anthropic's own evaluations at half the cost.

Honest caveats and risks worth knowing before you commit

Each vendor carries a real, current caveat that a purely benchmark-driven comparison would miss. Anthropic's Opus 5 ships stronger cybersecurity guardrails than Opus 4.8 and trails the company's own Mythos 5 on offensive-security tasks — for QA teams doing legitimate security testing, that means occasional refusals on benign work, and it is worth testing your actual prompts before standardising. Opus 5 also draws on a rate-limit pool separate from the Opus 4.x models, so migrating traffic neither frees nor inherits existing headroom. OpenAI's context-tier pricing on GPT-5.6 Sol is the kind of detail that shows up as a surprise invoice rather than a benchmark: agentic sessions cross 272K tokens routinely. DeepSeek's hosted app and API store consumer data on servers in mainland China, subject to Chinese data-access law, and it is restricted on government systems in several countries — a real consideration for regulated industries, even though the open-weight model can in principle be self-hosted elsewhere. Google's risk is roadmap rather than product: a flagship that has slipped three times, with its base model reportedly rebuilt, is a poor foundation for a multi-year commitment, and the Gemini CLI discontinuation in mid-2026 already broke CI automation for some teams. None of these makes a model unusable, but an honest comparison should surface them rather than only comparing headline scores.

FAQ

On vals.ai's independent SWE-bench Verified leaderboard as of late July 2026, Claude Opus 5 leads at 97.0%, with OpenAI's GPT-5.6 Sol at 96.2% and Claude Fable 5 at 95.0%. Opus 5 also tops Artificial Analysis's independent Intelligence Index at 61 (max effort) versus 59 for GPT-5.6 Sol. Those gaps are small enough that both are reasonable picks: choose Claude for flat 1M-context pricing and single-pass multi-file work, Codex for GitHub-native autonomous PR workflows, and factor in that Artificial Analysis measures Opus 5 as notably slow to first token.

Sources

Related guides