About
Every week a new AI model launches with a table of benchmark scores, and every leaderboard ranks them 1, 2, 3.
Most of those rankings are noise.
A 0.5-point gap on a benchmark with a few hundred questions is a coin flip. A model can lose a benchmark without tools and win it with them — same model, opposite verdict. Scores get reported by the labs themselves, on harnesses nobody else can reproduce.
The Model Gap reads each release and tells you one thing: which gaps are real, and which are just noise.
No composite score. No pretending 0.5 points means something.
The labels
- ✅ Real gap — outside the noise band, worth acting on
- ⚖️ Tie — statistically indistinguishable, ignore the ranking
- 🔧 Setup-dependent — the winner flips when the harness changes
- ⚠️ Unverified — vendor-reported only, no independent run yet
- 🚱 Tainted — the benchmark itself is saturated or contaminated
Get each verdict by email: themodelgap.substack.com · Contact: info@themodelgap.com