MATINTELLECT

👀 DeepSeek V4 Pro launched with the marketing line "we've closed the gap with the American frontier". The US NIST took the model for independent testing on closed datasets - and the public picture fell apart.

What CAISI (NIST's center for AI standards) found:
DeepSeek V4 Pro came out on April 24, 2026. On public benchmarks it looked convincing: SWE-Bench Verified - 80.6% vs 80.8% for Claude Opus 4.6, near parity. Codeforces rating 3,206 - the best result of any model at launch. Bloomberg wrote the same day that the model "hasn't closed the gap with the American leaders" - but that read like an editorial stance. CAISI added two datasets the model was guaranteed not to have seen in training: ARC-AGI-2 semi-private and its own PortBench. On those, the results broke down

Closed tests - the real gap:
🟢 Cybersecurity (CTF-Archive-Diamond, 285 tasks) - V4 Pro: 32% / GPT-5.5: 71%
🟢 Abstract reasoning (ARC-AGI-2, closed dataset) - V4 Pro: 46% / GPT-5.5: 79%
🟢 Software engineering (SWE-Bench Verified) - V4 Pro: 74% / GPT-5.5: 81%

CAISI's conclusion is clear-cut: the real lag behind the American frontier is about 8 months, not 3-6 like in DeepSeek's marketing. It's a steady trend since 2025 - the gap isn't shrinking. V4 does worst on long, complex tasks: where you need to hold many steps and a lot of context at once. That's exactly what separates a real agentic workload from synthetic tests. The mechanics are classic: the company picks its own benchmarks, takes the ones where its numbers peak, gets "almost caught up" headlines - NIST simply took tasks the model couldn't solve and wasn't trained to solve

Where DeepSeek V4 genuinely wins:
💡 Price - $0.145 per million tokens, 7x cheaper than GPT-5.5 and Claude Opus 4.7
💡 Open weights - you can deploy it locally and fine-tune it for your task
💡 Context - 1M tokens, MoE architecture (49B active out of 1.6T parameters)
💡 Code generation - Codeforces rating 3,206, a record at launch

On price vs capability, DeepSeek V4 Pro is still one of the best options on the market - especially for tasks where absolute quality isn't critical. The 8-month gap is real, but at 1/7 of the price it's a fair trade-off. The problem isn't the model - it's the "almost caught up" narrative built around it

⭐️ Yoshua Bengio, Turing Award laureate, chair of IASEAI:

"Independent evaluation of AI systems is not an option, it's a necessity. Companies' self-assessment can't be considered sufficient for understanding the real capabilities and risks of models"

💭 DeepSeek V4 is the strongest open Chinese model, and 7x cheaper than the frontier. That's a real argument for specific tasks. But "closed the gap with OpenAI" is about the press release, not the benchmark. Eight months of steady lag on closed tests - that's what the regulator sees, not the marketing team

Instagram | YouTube | Threads

Share:

No comments yet

Leave a Comment

Fields marked with an asterisk (*) are required