Skip to content
  • Models
  • Rankings
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

GPQA Diamond

GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. Each question is written by a subject-matter expert and designed so that even domain specialists need careful reasoning to identify the correct answer. We run the same fixed question set across provider endpoints to compare model capability, routing, and the practical cost of solving difficult scientific problems.

Last benchmark run Sep 5, 2026, 4:32 PM UTC

PaperGitHub
Model comparisonCost efficiencyLeaderboardExample problemsWhy we run itWhat scores tell youMethodologyAPI accessFAQ

Model comparison

Most Accurate

Favicon for google
Google: Gemini 3.1 Pro Preview

94.4%

Best Value

Favicon for google
Google: Gemini 3.7 Flash

$0.031/question

Fastest

Favicon for anthropic
Anthropic: Claude Fable 5.1

30s

Accuracy
Representative-run accuracy, best first.
Cost per question
Average cost per question, cheapest first.
Time per question
Average wall-clock time per question, fastest first.

Cost efficiency

Accuracy vs. cost (Pareto frontier)
One point per model, using default routing (not pinned to a provider) when available. The line is the Pareto frontier: no model beats these on both accuracy and cost.

Leaderboard

Top-level rows use default routing where available; click a row to expand provider-pinned results.

#ModelStd dev
1
Google: Gemini 3.1 Pro Preview
Pareto
94.4%--$0.202.1m16.7k
2
Google: Gemini 3.7 Flash
Pareto
94.3%--$0.03168s7.99k
3
OpenAI: GPT-5.5
93.8%±0.5pp$0.312.7m10.1k
4
OpenAI: GPT-5.6 Sol Pro
93.8%--$0.441.9m11.8k
5
SpaceXAI: Grok 4.6
93.3%--$0.126.6m20.1k
6
Google: Gemini 3.6 Flash
92.8%±0.4pp$0.06760s11.7k
7
Google: Gemini 3.5 Flash
92.8%±0.3pp$0.1480s15.6k
8
OpenAI: GPT-5.6 Sol
91.9%±0.5pp$0.07165s3.52k
9
MoonshotAI: Kimi K3
91.5%±1.1pp$0.156.5m10.6k
10
Anthropic: Claude Fable 5.1
90.9%--$0.1230s2.25k
11
MiniMax: MiniMax M3
Pareto
90.5%±2.1pp$0.0304.9m24.3k
12
OpenAI: GPT-5.4
90.3%±0.1pp$0.151.8m9.66k
13
OpenAI: GPT-5.6 Luna Pro
90.3%±0.4pp$0.0442.4m28.1k
14
Anthropic: Claude Opus 4.8
89.4%±1.5pp$0.201.6m7.73k
15
Amazon: Nova Micro 1.0
89.1%±0.7pp$0.221.8m8.52k
16
OpenAI: GPT-5.6 Terra
89.0%±0.6pp$0.03252s3.39k
17
Anthropic: Claude Opus 4.7
88.9%±0.8pp$0.221.8m8.66k
18
DeepSeek: DeepSeek V4 Pro 0813
88.9%±1.9pp$0.139.8m36.3k
19
Google: Gemini 3 Flash Preview
88.3%±0.1pp$0.134.2m43.9k
20
DeepSeek: DeepSeek V4 Flash Vision Exp
Pareto
88.2%±0.5pp$0.0234.7m22.8k
21
OpenAI: GPT-5.2
87.9%--$0.132.3m9.43k
22
OpenAI: GPT-5.6 Luna
Pareto
87.7%±0.7pp$0.00779s8.11k
23
Claude Opus 5
86.8%±2.2pp$0.1361s4.79k
24
Z.ai: GLM 5.3 Flash
Pareto
86.7%±1.5pp$0.0078.0m23.8k
25
DeepSeek: DeepSeek V4 Flash 0423
Pareto
86.6%±1.5pp$0.0044.7m16.5k
26
Anthropic: Claude Opus 4.5
86.6%±1.0pp$0.846.9m33.5k
27
Thinking Machines: Inkling Small
86.5%±1.1pp$0.0253.9m20.7k
28
DeepSeek: DeepSeek V4 Pro 0423
86.4%±1.5pp$0.0407.4m21.3k
29
OpenAI: GPT-5.1
86.2%±0.5pp$0.224.6m22.2k
30
OpenAI: GPT-5
86.0%--$0.214.5m20.7k
31
Qwen: Qwen3.5 397B A17B
85.9%±2.2pp$0.0374.8m12.6k
32
Z.ai: GLM 5.2
85.8%±2.7pp$0.0649.6m30.7k
33
Qwen: Qwen3.8 2.4T A95B
85.7%±1.3pp$0.0743.2m11.8k
34
Anthropic: Claude Opus 4.6
85.6%±0.9pp$0.748.1m29.3k
35
DeepSeek: DeepSeek V4 Flash 0731
85.2%±2.8pp$0.0098.2m24.5k
36
MiniMax: MiniMax M2.1
85.1%±1.3pp$0.0162.5m11.1k
37
MoonshotAI: Kimi K2.5
84.9%±3.1pp$0.07213.3m30.9k
38
Z.ai: GLM 5.3
84.8%±0.6pp$0.1416.2m31.1k
39
MiniMax: MiniMax M2.5
84.2%±2.2pp$0.0155.2m13.8k
40
Xiaomi: MiMo-V2.5-Pro
84.2%±1.0pp$0.0238.4m24.6k
41
MiniMax: MiniMax M2.7
84.1%±1.9pp$0.0256.5m20.1k
42
OpenAI: GPT-5.4 Mini
83.7%±0.7pp$0.0621.5m13.7k
43
MoonshotAI: Kimi K2 Thinking
83.6%±3.3pp$0.0446.8m17.7k
44
Qwen: Qwen3.5-122B-A10B
83.6%±3.6pp$0.0504.2m22k
45
Google: Gemini 3.5 Flash Lite
83.5%±0.0pp$0.03750s14.6k
46
Z.ai: GLM 4.7
83.5%±2.8pp$0.0457.9m22.7k
47
Google: Gemma 4 31B
83.4%±2.0pp$0.0045.3m9.61k
48
MoonshotAI: Kimi K2.6
83.3%±3.3pp$0.1111.8m34.3k
49
Thinking Machines: Inkling
83.2%±1.1pp$0.0975.7m23.8k
50
Z.ai: GLM 5.1
82.9%±2.6pp$0.119.0m31.4k
51
Anthropic: Claude Fable 5
82.8%±3.4pp$0.1743s3.13k
52
Anthropic: Claude Sonnet 4.5
82.6%±1.6pp$0.324.8m20.9k
53
Anthropic: Claude Sonnet 5
82.2%±2.8pp$0.162.7m15.9k
54
NVIDIA: Nemotron 3 Ultra
82.2%±2.3pp$0.0834.0m23.8k
55
DeepSeek: DeepSeek V3.2
82.0%±2.1pp$0.0045.1m10.4k
56
Qwen: Qwen3.8 27B
81.8%±1.7pp$0.0363.8m13.4k
57
Qwen: Qwen3.5-35B-A3B
81.7%±3.7pp$0.0183.8m16.5k
58
MoonshotAI: Kimi K2.7 Code
80.7%±18.0pp$0.0777.5m21.2k
59
Xiaomi: MiMo-V2-Flash
Pareto
80.7%±4.7pp$0.00384s10.6k
60
Qwen: Qwen3.6 35B A3B
80.6%±4.4pp$0.0375.8m35.8k
61
OpenAI: GPT-5.3 Chat
80.5%--$0.02424s1.63k
62
DeepSeek: DeepSeek V3.1 Terminus
80.4%±2.7pp$0.0094.7m9.61k
63
Meta: Muse Glimmer 30B
80.2%±1.8pp$0.0224.5m17.5k
64
Google: Gemini 3.1 Flash Lite
80.1%±0.5pp$0.04281s27.6k
65
Qwen: Qwen3 235B A22B Thinking 2507
80.0%±3.3pp$0.0295.9m14.7k
66
Google: Gemini 2.5 Pro
79.9%±1.1pp$0.253.8m25.4k
67
OpenAI: GPT-5.2 Chat
79.8%--$0.03633s2.48k
68
OpenAI: GPT-5 Mini
79.6%±0.8pp$0.0444.2m22.1k
69
Qwen: Qwen3.6 27B
79.6%±4.5pp$0.07110.7m27.8k
70
DeepSeek: DeepSeek V3.1
79.3%±3.5pp$0.0105.1m9.34k
71
DeepSeek: R1 0528
79.1%±1.4pp$0.03410.4m15.6k
72
Z.ai: GLM 4.6
78.8%±4.3pp$0.0369.7m18.5k
73
DeepSeek: DeepSeek V3.2 Exp
78.8%±3.1pp$0.0047.8m10.3k
74
OpenAI: GPT-5.4 Nano
77.9%±0.2pp$0.01370s10.4k
75
Qwen: Qwen3.5-9B
77.7%±2.6pp$0.0059.7m32.1k
76
Z.ai: GLM 5
77.5%±7.0pp$0.0688.7m30.1k
77
Anthropic: Claude Sonnet 4
76.7%±1.9pp$0.273.0m17.6k
78
StepFun: Step 3.7 Flash
76.6%--$0.0807.1m69.8k
79
Xiaomi: MiMo-V2.5
76.3%±5.6pp$0.0109.8m29.2k
80
MoonshotAI: Kimi K2 0905
75.9%±1.7pp$0.0101.9m3.6k
81
Mistral: Mistral Small 4
75.8%--$0.0152.9m24k
82
OpenAI: gpt-oss-120b
75.3%±2.9pp$0.0072.9m14k
83
Auto Router (Beta)
74.7%--$0.0614.3m31.8k
84
Qwen: Qwen3 235B A22B Instruct 2507
74.4%±2.1pp$0.0042.7m5.6k
85
Qwen: Qwen3 Coder Next
74.1%±2.1pp$0.00973s9.27k
86
Google: Gemma 4 26B A4B
73.9%±4.0pp$0.0087.6m21.8k
87
inclusionAI: Ling 3.0 Flash
Pareto
73.0%--$0.0033.3m31.4k
88
Google: Gemini 2.5 Flash
72.7%±0.5pp$0.0592.0m23.5k
89
Anthropic: Claude Haiku 4.5
72.1%±0.5pp$0.194.0m38k
90
OpenAI: GPT-5 Nano
70.9%--$0.0133.7m33.1k
91
Qwen: Qwen3 Next 80B A3B Instruct
70.7%±1.9pp$0.00767s6.42k
92
Qwen: Qwen3 VL 235B A22B Instruct
70.5%±1.8pp$0.0062.7m4.32k
93
NVIDIA: Nemotron 3.5 Lightning
69.0%±0.6pp$0.0115.2m43.9k
94
Gemma 4 26B A4B IT (free)
68.5%----6.4m6.99k
95
OpenAI: gpt-oss-20b
66.6%±2.6pp$0.0066.4m33.7k
96
Meta: Llama 4 Maverick
Pareto
65.9%±1.6pp$0.00259s2.2k
97
OpenAI: GPT-4.1
64.9%±1.6pp$0.01614s1.77k
98
Qwen: Qwen3 VL 30B A3B Instruct
64.8%±2.1pp$0.0063.0m9.55k
99
Z.ai: GLM 4.5 Air
64.5%±6.5pp$0.0135.4m14.4k
100
OpenAI: GPT-4.1 Mini
64.1%±0.2pp$0.00431s2.32k
101
Qwen: Qwen3 30B A3B Instruct 2507
Pareto
63.9%±2.0pp$0.00172s4.75k
102
Qwen: Qwen3 Coder 480B A35B
61.7%±3.2pp$0.00232s1.22k
103
Qwen: Qwen3 30B A3B
61.4%±2.1pp$0.0042.6m7.67k
104
Qwen: Qwen3 32B
61.4%±2.0pp$0.0043.1m10.3k
105
DeepSeek: DeepSeek V3
60.9%--$0.00274s2.21k
106
NVIDIA: Nemotron 3 Nano 30B A3B
60.7%±2.7pp$0.0127.4m58.4k
107
Qwen: Qwen3 14B
59.4%--$0.0038.5m11k
108
DeepSeek: DeepSeek V3 0324
58.1%±11.3pp$0.00259s1.8k
109
Google: Gemini 2.5 Flash Lite
55.0%±0.8pp$0.0172.1m42.5k
110
Anthropic: Claude Sonnet 4.6
54.4%±0.6pp$0.6110.5m40.4k
111
Qwen: Qwen3 Coder 30B A3B Instruct
52.0%±0.2pp$0.0041.5m2.72k
112
Qwen: Qwen3 VL 8B Instruct
51.8%±1.8pp$0.0042.0m8.03k
113
OpenAI: GPT-4o (2024-08-06)
51.7%--$0.01817s1.63k
114
Z.ai: GLM 4.7 Flash
51.4%±3.9pp$0.0086.0m21.3k
115
OpenAI: GPT-4.1 Nano
Pareto
50.9%±0.9pp$0.0008220s1.86k
116
OpenAI: GPT-4o (2024-05-13)
50.7%--$0.02411s1.32k
117
OpenAI: GPT-4o
50.3%±0.8pp$0.01313s1.13k
118
Meta: Llama 3.3 70B Instruct
Pareto
49.5%±2.6pp$0.0007638s1.35k
119
Qwen2.5 72B Instruct
44.6%--$0.000981.5m1.76k
120
Qwen: Qwen2.5 VL 72B Instruct
44.4%±1.5pp$0.00259s1.59k
121
OpenAI: GPT-4o-mini
43.4%±0.9pp$0.00121s1.49k
122
Qwen: Qwen2.5 7B Instruct
Pareto
32.8%--$0.0005336s2.04k
123
Mistral: Mistral Nemo
Pareto
31.6%±2.1pp$0.0000848s363
124
Meta: Llama 3.1 8B Instruct
28.6%±3.0pp$0.0006252s7.82k
125
Sao10K: Llama 3 8B Lunaris
Pareto
26.9%±0.7pp$0.0000639s592
126
inclusionAI: Ling 3.0 Flash Fin
19.7%--$0.01941.6m106k
127
Meta: Llama 3.2 3B Instruct
11.6%--$0.0001530s1.84k
128
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
10.3%--$0.01822s11.8k

Example problems

GPQA uses four-choice questions that require more than recalling a definition. These representative examples show the format and the range of scientific domains without reproducing items from the benchmark's protected question pool.

Biology

A researcher observes that a membrane protein is synthesized on ribosomes attached to the rough endoplasmic reticulum. Which destination is most consistent with this protein entering the secretory pathway?

  1. A.The cytosol, where it remains soluble
  2. B.The nucleus, after import through a nuclear pore
  3. C.A membrane of the endomembrane system or the cell surface
  4. D.The mitochondrial matrix through a TOM/TIM complex

Answer: C

Ribosomes on the rough ER synthesize proteins destined for secretion or insertion into the endomembrane system, including the plasma membrane.

Physics

A spacecraft is far from other bodies and fires its engine in the direction opposite to its velocity. Ignoring mass loss during the brief burn, what happens immediately to its speed?

  1. A.It increases because the exhaust carries away backward momentum
  2. B.It decreases because the thrust points opposite to its velocity
  3. C.It remains unchanged because thrust only changes direction
  4. D.It becomes zero because the spacecraft is in free space

Answer: B

An impulse opposite the velocity vector reduces the spacecraft’s momentum and therefore its speed during the burn.

Chemistry

Why does adding a small amount of a common ion generally reduce the solubility of a sparingly soluble ionic solid in water?

  1. A.The common ion increases the solid’s lattice energy
  2. B.The common ion shifts the dissolution equilibrium toward the solid
  3. C.The common ion converts every dissolved ion into a neutral molecule
  4. D.The common ion removes solvent molecules from the solution

Answer: B

The added ion raises the concentration of a dissolution product, so Le Chatelier’s principle shifts the equilibrium toward the undissolved solid.

Why we run this benchmark

GPQA is a broad graduate-level reasoning test across biology, physics, and chemistry, so it gives us a cheap, high-floor signal that a deployment is healthy. A model that normally clears these questions but suddenly drops usually points to something broken in the endpoint or routing rather than the questions themselves.

Because we run the same fixed question set across provider endpoints, a large accuracy gap between providers serving the same model is a quick way to catch a misconfigured or degraded endpoint. The cost and latency columns show what that reasoning quality costs to serve.

What the scores can and can't tell you

GPQA is a narrow, high-difficulty evaluation, not a complete measure of general intelligence or usefulness. A score reflects performance on expert-written multiple-choice science questions and should be considered alongside coding, instruction-following, factuality, and other evaluations.

Scores can be sensitive to sampling settings, answer-position handling, and the number of repeated runs. Small differences may not be meaningful when models have similar sample counts, so the leaderboard includes run variability and cost context rather than presenting accuracy alone.

The benchmark is publicly described, and some questions may eventually appear in training data. We avoid reproducing the private question pool here, but no public benchmark can guarantee that every future evaluation item is uncontaminated.

Methodology

Scores aggregate successful runs, weighted by question count, with a minimum sample threshold per model-provider pair. A model's headline score uses default routing when available; otherwise it falls back to the median provider. Cost, time, and output-token figures are per-question averages from the same runs. Best value is the cheapest Pareto-optimal model within five percentage points of the top score.

GPQA Diamond is described in the original paper. See the docs for routing details, or browse all models to try one.

API access

These scores are available through OpenRouter's public benchmarks API, so you can retrieve the same model-level results programmatically.

GET https://openrouter.ai/api/v1/benchmarks?source=openrouter
Authorization: Bearer <API key>

Use task_type=intelligence to filter to gpqa_diamond. Each item represents one model and includes accuracy, accuracy_stddev, avg_cost_per_task, total_tasks, and last_run_timestamp. See the benchmarks API docs.

Frequently asked questions

GPQA Diamond is a graduate-level multiple-choice benchmark in biology, physics, and chemistry. Each question is written by a subject-matter expert and designed so that even domain specialists need careful reasoning to identify the correct answer.
Every model answers the same fixed question set through real provider endpoints, and results are aggregated across repeated runs. A model’s headline score is a single representative result rather than its best-performing provider, and the cost, time, and output-token figures are per-question averages from those same runs.
Every run goes to a real provider endpoint, so provider behavior is part of the measurement. A large accuracy gap between providers serving the same model usually points to a misconfigured or degraded endpoint rather than to the questions.
Yes. The public benchmarks API returns the same model-level results from GET https://openrouter.ai/api/v1/benchmarks?source=openrouter with an API key. Filter with task_type=intelligence to reach gpqa_diamond.
It is a narrow, high-difficulty evaluation rather than a measure of general usefulness. Scores are sensitive to sampling settings, answer-position handling, and the number of repeated runs, so small differences between models with similar sample counts may not be meaningful.