DeepSeek
DeepSeek

🏆 Benchmarks

DeepSeek V4 Pro vs GPT-4o vs Claude — MMLU 89.2%, Codeforces 3206, MATH-500 85.3%. Leading in code, math and reasoning across 6 benchmarks.

📊 Evaluation Summary

Leading across code, math and reasoning domains.

6

Total Benchmarks

6

Leading

89.1%

Average Score

Beats GPT-4o

Select Model

Category

Multi-Dimensional Capability Radar

DeepSeek V4 Pro six-dimension capability profile.

ReasoningCodeMathKnowledgeSTEMAgent
Reasoning: 90.1Code: 93.5Math: 85.3Knowledge: 89.2STEM: 88Agent: 85.4

Detailed Comparison

MMLU

Multitask language understanding across 57 subjects.
🥇 DeepSeek V4 Pro
89.2%
DeepSeek V4 Flash
85.1%
🥉 GPT-4o
86.4%
🥈 Claude
88.7%

GPQA Diamond

Graduate-level expert reasoning questions.
🥇 DeepSeek V4 Pro
90.1%
DeepSeek V4 Flash
83.7%
🥉 GPT-4o
85.2%
🥈 Claude
87.3%

HumanEval

Python code generation from docstrings.
🥇 DeepSeek V4 Pro
87.5%
DeepSeek V4 Flash
82.3%
🥉 GPT-4o
84.1%
🥈 Claude
86.2%

MATH-500

Competition-level mathematics problems.
🥇 DeepSeek V4 Pro
85.3%
🥉 DeepSeek V4 Flash
78.9%
GPT-4o
76.6%
🥈 Claude
80.1%

Codeforces

Competitive programming Elo rating.
🥇 DeepSeek V4 Pro
3206
🥉 DeepSeek V4 Flash
2850
GPT-4o
2800
🥈 Claude
2950

LiveCodeBench

Contamination-free live coding benchmark.
🥇 DeepSeek V4 Pro
93.5%
🥉 DeepSeek V4 Flash
89.2%
GPT-4o
88.2%
🥈 Claude
90.1%

Full Data Table

Benchmark DeepSeek V4 ProDeepSeek V4 FlashGPT-4oClaude
MMLU 89.2% 🥇 85.1% 86.4% 88.7%
GPQA Diamond 90.1% 🥇 83.7% 85.2% 87.3%
HumanEval 87.5% 🥇 82.3% 84.1% 86.2%
MATH-500 85.3% 🥇 78.9% 76.6% 80.1%
Codeforces 3206 🥇 2850 2800 2950
LiveCodeBench 93.5% 🥇 89.2% 88.2% 90.1%

📈 Performance Trend

Improvement from V3 to V4 series.

Code +12% Reasoning +8% Math +15%

Methodology

Data Source
DeepSeek official technical report
Evaluation Date
2026-06-30
Sampling Settings
temperature=0, top_p=1