GPQA Benchmark: GPQA Diamond Explained, Scores by Model, and How to Run It

GPQA is a 448-question graduate-level science benchmark (biology, physics, chemistry) built to be Google-proof. GPQA Diamond is its 198-question hardest subset: PhD experts score 81.3%, skilled non-experts 22.1%, chance is 25%. 2026 model scores, why reports disagree, and how to run it.

October 5, 2026 · 2 min read

What Is the GPQA Benchmark?

Last updated October 5, 2026.

GPQA (Graduate-Level Google-Proof Q&A) is a multiple-choice benchmark of 448 questions in biology, physics, and chemistry, written and validated by people who hold or are pursuing PhDs in those fields. David Rein and co-authors released it in November 2023 (arXiv 2311.12022). Experts reach 65% on the main set, or 74% after discounting mistakes they acknowledged in retrospect. Highly skilled non-experts reach 34% after spending more than 30 minutes per question with unrestricted web access, which is where the Google-proof name comes from. The strongest GPT-4 baseline in the paper scored 39%. GPQA Diamond, the 198-question hardest subset, is the number every 2026 model card reports, and frontier models now land between 91% and 94% on it.

198
questions in GPQA Diamond
25%
random-guess baseline (4 options)
81.3%
expert validator accuracy on Diamond
96.3%
top independent score (Artificial Analysis)

Each question has four answer choices. The median question, including its choices, is 561 characters, or 146 tokens. Question writers were paid bonuses when experts agreed with their answer and non-experts got it wrong, so the collection process selected for questions that reward domain knowledge over search skill. Non-expert validators were experts in a different field: a physicist validating a biology question, for example.

GPQA Diamond vs GPQA Main vs GPQA Extended

The authors collected 564 questions and held back 18 as an unreleased test set, leaving 546. Those 546 are GPQA Extended. The two curated subsets filter on how the validators did.

GPQA splits (Rein et al. 2023, Table 2)
SplitQuestionsSelection ruleExpert accuracyNon-expert accuracy
GPQA Extended546All released questions64.8%34.1%
GPQA (main)448At least 1 of 2 experts agree, at most 2 of 3 non-experts correct71.9%*30.4%*
GPQA Diamond198Both experts agree, at most 1 of 3 non-experts correct81.3%*22.1%*

* The paper flags main and Diamond validator accuracies as skewed by selection effects, since the subsets were chosen using those same validator answers.

When a model card says "GPQA" with no qualifier in 2026, it almost always means Diamond. The original paper reported its GPT-4 baselines on the main set, so the 39% figure and today's Diamond scores are measured on different question pools.

GPQA Human Baselines: Experts, Non-Experts, and Chance

Reference points for reading a GPQA score
BaselineScoreSetSource
Random guessing25%Any split4 answer choices
Skilled non-experts, web access34%MainRein et al. 2023
GPT-4 (best paper baseline)39%MainRein et al. 2023
Domain experts65% (74% adjusted)MainRein et al. 2023
PhD experts recruited by OpenAI69.7%DiamondOpenAI, via Epoch AI

Expertise varied by subfield. On the extended set, expert validators scored 81% on organic chemistry questions, against 55% for the rest of chemistry. A 90% model score on Diamond beats a typical expert validator answering across a whole domain. A panel of specialists, each answering only inside their own subfield, would likely score higher than the single-validator figures in the paper.

GPQA Diamond Scores by Model (2026)

Most published GPQA Diamond numbers are vendor-reported in model cards. The table groups them by the card that reported them, because comparisons inside one card share a harness and comparisons across cards do not.

Independent: Artificial Analysis GPQA Diamond leaderboard (top 3 of 615 models)
ModelGPQA Diamond
GPT-6 Astra (xhigh)96.3%
GPT-6 Astra (max)96.1%
Gemini 3.8 Flash (High)95.3%
Reported in the Kimi K3 model card (Moonshot AI)
ModelGPQA DiamondServed on Morph
GPT-5.6 Sol (max)94.1No
Kimi K3 (max)93.5morph-kimik3
GPT-5.5 (xhigh)93.5No
Claude Fable 5 (max)92.6No
GLM-5.2 (max)91.2morph-glm53-744b
Claude Opus 4.8 (max)91.0No
Reported in the GLM-5.2 model card (Z.ai)
ModelGPQA DiamondServed on Morph
Gemini 3.1 Pro94.3No
Claude Opus 4.893.6No
GPT-5.593.6No
MiniMax M393morph-minimax3-428b
GLM-5.291.2morph-glm53-744b
DeepSeek-V4-Pro90.1No
Qwen3.7-Max90No
GLM-5.186.2No
Reported in the Qwen3.6-27B model card (Alibaba Qwen)
ModelGPQA DiamondServed on Morph
Qwen3.5-397B-A17B88.4morph-qwen35-397b
Qwen3.6-27B87.8morph-qwen36-27b
Claude 4.5 Opus87.0No
Qwen3.6-35B-A3B86.0No
Qwen3.5-27B85.5No
Gemma4-31B84.3morph-gemma4-31b

The open-weight gap on GPQA Diamond is small. Kimi K3 and MiniMax M3 report within a point of GPT-5.5 and Claude Opus 4.8 in their respective cards, and a 27B dense model (Qwen3.6-27B) reports 87.8, above the 87.0 its card lists for Claude 4.5 Opus. The full benchmark tables for each model live on the Kimi K3, MiniMax M3, GLM-5.2, and Qwen 3.6 pages, with Anthropic models on Claude benchmarks.

Why GPQA Diamond Scores Disagree Between Reports

±4.2 pts
“95% confidence interval for a single GPQA Diamond run at 90% accuracy, from 198 questions.”

One question is 1/198 of the score, or about 0.5 points. At 90% accuracy the standard error of a single run is the square root of 0.9 × 0.1 / 198, about 2.1 points, so the 95% interval spans roughly 4.2 points either side. Two models 2 points apart on one run each cannot be ranked from that evidence alone.

Harness choices add a second source of spread. Claude Opus 4.8 appears at 93.6 in the GLM-5.2 card and at 91.0 (max effort) in the Kimi K3 card. GPT-5.5 appears at 93.6 and 93.5, which agree. The differences come from reasoning effort, temperature and top-p (Moonshot reports temperature 1.0 and top-p 0.95 for GPQA), prompt wording, and how the final letter is parsed out of a long reasoning trace. Epoch AI uses the simple-evals prompt, which asks the model to end with a line of the form "ANSWER: LETTER".

Reading a GPQA claim

Compare models inside one model card or one independent leaderboard. Treat gaps under 4 points from single runs as ties. Check the reasoning-effort setting, since the same model at max and at default effort can differ by more than the gap to its nearest competitor.

Is the GPQA Benchmark Saturated?

For separating frontier models, mostly yes. The top independent score (96.3%) is above the 81.3% expert-validator accuracy on Diamond and well above the 69.7% that OpenAI's recruited PhD panel scored. The paper itself reports expert accuracy rising from 65% to 74% once experts' acknowledged mistakes are discounted, so some residual disagreement over answers is part of the dataset. The last few points sit inside both that disagreement and the 4-point noise band.

GPQA Diamond still works as a floor check. A model under 80% in 2026 is well behind the open-weight frontier on graduate-level science recall and reasoning, and a quick Diamond run catches quantization or serving regressions, since it is short, cheap, and has a fixed answer key. Contamination is the other caveat: the dataset ships with a canary string and a password-protected archive to keep it out of training crawls, but nothing guarantees every lab filtered it.

How to Run GPQA Diamond (lm-evaluation-harness)

The dataset is gated on Hugging Face at Idavidrein/gpqa; accept the terms there and log in with huggingface-cli login before running. EleutherAI's lm-evaluation-harness ships every split in zero-shot, n-shot, generative, and chain-of-thought variants (gpqa_diamond_zeroshot, gpqa_diamond_cot_zeroshot, gpqa_main_n_shot, and so on). For chat models served behind an OpenAI-compatible API, the chain-of-thought generative task is the closest match to how labs report the number.

pip install "lm_eval[api]"
huggingface-cli login

export OPENAI_API_KEY=YOUR_MORPH_API_KEY

lm_eval --model local-chat-completions \
  --model_args model=morph-kimik3,base_url=https://api.morphllm.com/v1/chat/completions,num_concurrent=8,max_retries=3 \
  --tasks gpqa_diamond_cot_zeroshot \
  --apply_chat_template \
  --log_samples \
  --output_path results/

Reasoning models write long traces before the answer letter. If the generation limit cuts the trace off, the item scores as wrong, so raise the limit before blaming the model. Keep the per-sample logs: a GPQA number without them cannot be checked for answer-parsing failures.

Cost is dominated by output tokens. At an assumed 10,000 output tokens per question, one Diamond pass on morph-kimik3 ($14.00/M output) costs about $27.72. Run it three to five times and average if you need to separate models closer than 4 points.

GPQA vs SWE-bench Pro, Terminal-Bench, and Other LLM Benchmarks

Where GPQA fits
BenchmarkWhat it measuresFormatMorph guide
GPQA DiamondGraduate-level science (bio, physics, chem)198 four-option questionsThis page
SWE-bench ProResolving real repository issuesAgentic patch generationSWE-bench Pro guide
Terminal-BenchCompleting tasks in a shellAgentic, containerizedClaude benchmarks
lm-eval-harness tasksMMLU-Pro, BBH, MATH, IFEval, GPQAAcademic multiple choice and generationEval harness guide

GPQA tells you how much graduate science a model can recall and reason through in a single answer. It says little about coding agents. For that, read SWE-bench Pro and the coding sections of the Claude benchmarks page, and for picking an open model to self-host, the open-source LLM guide.

GPQA Benchmark FAQ

What is the GPQA benchmark?

GPQA (Graduate-Level Google-Proof Q&A) is a multiple-choice benchmark of 448 questions in biology, physics, and chemistry, written by domain experts and released by Rein et al. in November 2023 (arXiv 2311.12022). PhD-level experts score 65% on it (74% after discounting acknowledged mistakes), while skilled non-experts with unrestricted web access score 34%.

What is GPQA Diamond?

GPQA Diamond is the 198-question subset of GPQA where both expert validators answered correctly and the majority of non-expert validators answered incorrectly. The paper reports 81.3% expert accuracy and 22.1% non-expert accuracy on Diamond, though both are skewed by the selection rule. It is the split model cards and leaderboards report.

What is a good GPQA Diamond score?

Random guessing scores 25%. GPT-4 scored 39% on the main set in the original paper, and a panel of PhD experts recruited by OpenAI scored 69.7% on Diamond. Frontier models in 2026 report 91% to 94%, and the top model measured independently by Artificial Analysis scores 96.3%. A gap under about 4 points between two single runs is within statistical noise.

Which model has the highest GPQA Diamond score?

On the Artificial Analysis independent leaderboard (615 models evaluated), GPT-6 Astra at xhigh reasoning leads at 96.3%, followed by GPT-6 Astra at max reasoning at 96.1% and Gemini 3.8 Flash (High) at 95.3%. Among open-weight models, Kimi K3 reports 93.5 and MiniMax M3 is listed at 93 in the GLM-5.2 model card.

Why do different sources report different GPQA scores for the same model?

Prompt format, reasoning effort, sampling temperature, answer extraction, and the number of runs averaged all differ between labs. Claude Opus 4.8 is listed at 93.6 in the GLM-5.2 model card and at 91.0 in the Kimi K3 model card. With only 198 questions, that 2.6-point gap is about five questions.

How do I run GPQA Diamond myself?

Accept the terms on the gated Hugging Face dataset Idavidrein/gpqa, log in with huggingface-cli, and run EleutherAI's lm-evaluation-harness with a task such as gpqa_diamond_cot_zeroshot. Any OpenAI-compatible endpoint works through the local-chat-completions model type, including Morph's api.morphllm.com/v1.

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.

Talk to us about a private deployment

Run GPQA Diamond against open models on Morph

Kimi K3, MiniMax M3, GLM, and Qwen 3.6 behind one OpenAI-compatible endpoint. Point lm-evaluation-harness at api.morphllm.com/v1 and benchmark them on your own harness settings.

Sources