🏆 Benchmarks
DeepSeek V4 Pro vs GPT-4o vs Claude — MMLU 89.2%, Codeforces 3206, MATH-500 85.3%. Leading in code, math and reasoning across 6 benchmarks.
📊 Evaluation Summary
Leading across code, math and reasoning domains.
6
Total Benchmarks
6
Leading
89.1%
Average Score
✓
Beats GPT-4o
Select Model
Category
Multi-Dimensional Capability Radar
DeepSeek V4 Pro six-dimension capability profile.
Reasoning: 90.1Code: 93.5Math: 85.3Knowledge: 89.2STEM: 88Agent: 85.4
Detailed Comparison
MMLU
Multitask language understanding across 57 subjects. 🥇 DeepSeek V4 Pro 89.2%
DeepSeek V4 Flash 85.1%
🥉 GPT-4o 86.4%
🥈 Claude 88.7%
GPQA Diamond
Graduate-level expert reasoning questions. 🥇 DeepSeek V4 Pro 90.1%
DeepSeek V4 Flash 83.7%
🥉 GPT-4o 85.2%
🥈 Claude 87.3%
HumanEval
Python code generation from docstrings. 🥇 DeepSeek V4 Pro 87.5%
DeepSeek V4 Flash 82.3%
🥉 GPT-4o 84.1%
🥈 Claude 86.2%
MATH-500
Competition-level mathematics problems. 🥇 DeepSeek V4 Pro 85.3%
🥉 DeepSeek V4 Flash 78.9%
GPT-4o 76.6%
🥈 Claude 80.1%
Codeforces
Competitive programming Elo rating. 🥇 DeepSeek V4 Pro 3206
🥉 DeepSeek V4 Flash 2850
GPT-4o 2800
🥈 Claude 2950
LiveCodeBench
Contamination-free live coding benchmark. 🥇 DeepSeek V4 Pro 93.5%
🥉 DeepSeek V4 Flash 89.2%
GPT-4o 88.2%
🥈 Claude 90.1%
Full Data Table
| Benchmark | DeepSeek V4 Pro | DeepSeek V4 Flash | GPT-4o | Claude |
|---|---|---|---|---|
| MMLU | 89.2% 🥇 | 85.1% | 86.4% | 88.7% |
| GPQA Diamond | 90.1% 🥇 | 83.7% | 85.2% | 87.3% |
| HumanEval | 87.5% 🥇 | 82.3% | 84.1% | 86.2% |
| MATH-500 | 85.3% 🥇 | 78.9% | 76.6% | 80.1% |
| Codeforces | 3206 🥇 | 2850 | 2800 | 2950 |
| LiveCodeBench | 93.5% 🥇 | 89.2% | 88.2% | 90.1% |
📈 Performance Trend
Improvement from V3 to V4 series.
Code +12% Reasoning +8% Math +15%
Methodology
- Data Source
- DeepSeek official technical report
- Evaluation Date
- 2026-06-30
- Sampling Settings
temperature=0, top_p=1