The Model Gap

DeepSeek V4 Pro: Real or Noise?

August 17, 2026 · Verdict #1

On August 13, DeepSeek released V4 Pro with the now-standard launch artifact: a chart of agentic benchmarks where it trades wins with Claude Opus 4.8.

The chart is real data. It is also doing what launch charts always do: presenting every number at face value, as if they were all equally meaningful.

They are not. Here is the same chart, read honestly.

First, the provenance

Every number below comes from DeepSeek's own launch material. Vendor-reported, vendor-chosen benchmarks, vendor-chosen harness, no independent replication yet. That doesn't make the numbers false — it makes them ⚠️ unverified until third parties rerun them. Keep that discount rate in mind for everything that follows.

The headline flip: HLE, with and without tools

The most interesting row is Humanity's Last Exam, reported two ways: without tools and with tools.

SettingDeepSeek V4 ProClaude Opus 4.8Verdict
HLE, no tools42.749.8Opus by 7.1
HLE, with tools60.057.9V4 Pro by 2.1

Same two models. Same benchmark. Opposite verdicts. The only thing that changed is the scaffolding.

Look at the deltas instead of the rankings: give V4 Pro tools and it gains +17.3 points. Give Opus the same tools and it gains +8.1. That asymmetry is the actual finding — DeepSeek clearly trained hard on tool use, and it shows.

But it also means "which model is smarter" is the wrong question here. 🔧 Setup-dependent. If your workload gives models tools — agents, search, code execution — V4 Pro's number is the one that describes your world. If it doesn't, Opus is still ahead, by a real margin.

(A quiet footnote the headlines skipped: the rightmost column of DeepSeek's own chart, Fable 5, posts 53.3 and 63.0 on this same row — ahead in both settings. Launch charts pick their duels.)

The noise: Agents' Last Exam

V4 Pro scores 25.2. Opus 4.8 scores 25.7.

That is a 0.5-point gap. On benchmarks in this size class, run-to-run variance alone moves scores by more than that. Rank these two models on this row and you are ranking dice.

⚖️ Tie. Anyone telling you otherwise is selling precision that does not exist.

The real story: DeepSWE

V4 Pro: 62.7. Opus 4.8: 58.0. A 4.7-point lead for DeepSeek — plausibly meaningful, pending replication.

But the solider number sits inside DeepSeek's own column: the previous release, V4 Pro Preview, scored 12.8 on this row. That is a jump of roughly fifty points in one generation, measured by the same vendor on the same harness both times — no cross-lab excuse available.

✅ Real — not the duel with Opus, which needs a third party to confirm, but the trajectory. Whatever DeepSeek changed in its agentic software-engineering training between Preview and Pro, it worked.

For balance: on Toolathlon-Verified, Opus stays ahead 76.2 to 74.1 — a 2.1-point edge that is itself borderline noise. ⚖️ Call it a tie until someone reruns it.

The one thing to remember

DeepSeek V4 Pro genuinely caught up with frontier agents when given tools, and its generation-over-generation jump is real and steep. The "beats Opus 4.8" headline, though, lives inside the margin of error and inside one particular harness. If you run tool-using agents, add it to your test list. If you were about to pick a model off this chart's rankings — don't.


All numbers are from DeepSeek's launch chart (vendor-reported). No independent replications existed at time of writing; we'll revisit when they land. Labels: ✅ real gap · ⚖️ tie · 🔧 setup-dependent · ⚠️ unverified.