What Is the GPQA Benchmark?
Last updated October 5, 2026.
GPQA (Graduate-Level Google-Proof Q&A) is a multiple-choice benchmark of 448 questions in biology, physics, and chemistry, written and validated by people who hold or are pursuing PhDs in those fields. David Rein and co-authors released it in November 2023 (arXiv 2311.12022). Experts reach 65% on the main set, or 74% after discounting mistakes they acknowledged in retrospect. Highly skilled non-experts reach 34% after spending more than 30 minutes per question with unrestricted web access, which is where the Google-proof name comes from. The strongest GPT-4 baseline in the paper scored 39%. GPQA Diamond, the 198-question hardest subset, is the number every 2026 model card reports, and frontier models now land between 91% and 94% on it.
Each question has four answer choices. The median question, including its choices, is 561 characters, or 146 tokens. Question writers were paid bonuses when experts agreed with their answer and non-experts got it wrong, so the collection process selected for questions that reward domain knowledge over search skill. Non-expert validators were experts in a different field: a physicist validating a biology question, for example.
GPQA Diamond vs GPQA Main vs GPQA Extended
The authors collected 564 questions and held back 18 as an unreleased test set, leaving 546. Those 546 are GPQA Extended. The two curated subsets filter on how the validators did.
| Split | Questions | Selection rule | Expert accuracy | Non-expert accuracy |
|---|---|---|---|---|
| GPQA Extended | 546 | All released questions | 64.8% | 34.1% |
| GPQA (main) | 448 | At least 1 of 2 experts agree, at most 2 of 3 non-experts correct | 71.9%* | 30.4%* |
| GPQA Diamond | 198 | Both experts agree, at most 1 of 3 non-experts correct | 81.3%* | 22.1%* |
* The paper flags main and Diamond validator accuracies as skewed by selection effects, since the subsets were chosen using those same validator answers.
When a model card says "GPQA" with no qualifier in 2026, it almost always means Diamond. The original paper reported its GPT-4 baselines on the main set, so the 39% figure and today's Diamond scores are measured on different question pools.
GPQA Human Baselines: Experts, Non-Experts, and Chance
| Baseline | Score | Set | Source |
|---|---|---|---|
| Random guessing | 25% | Any split | 4 answer choices |
| Skilled non-experts, web access | 34% | Main | Rein et al. 2023 |
| GPT-4 (best paper baseline) | 39% | Main | Rein et al. 2023 |
| Domain experts | 65% (74% adjusted) | Main | Rein et al. 2023 |
| PhD experts recruited by OpenAI | 69.7% | Diamond | OpenAI, via Epoch AI |
Expertise varied by subfield. On the extended set, expert validators scored 81% on organic chemistry questions, against 55% for the rest of chemistry. A 90% model score on Diamond beats a typical expert validator answering across a whole domain. A panel of specialists, each answering only inside their own subfield, would likely score higher than the single-validator figures in the paper.
GPQA Diamond Scores by Model (2026)
Most published GPQA Diamond numbers are vendor-reported in model cards. The table groups them by the card that reported them, because comparisons inside one card share a harness and comparisons across cards do not.
| Model | GPQA Diamond |
|---|---|
| GPT-6 Astra (xhigh) | 96.3% |
| GPT-6 Astra (max) | 96.1% |
| Gemini 3.8 Flash (High) | 95.3% |
| Model | GPQA Diamond | Served on Morph |
|---|---|---|
| GPT-5.6 Sol (max) | 94.1 | No |
| Kimi K3 (max) | 93.5 | morph-kimik3 |
| GPT-5.5 (xhigh) | 93.5 | No |
| Claude Fable 5 (max) | 92.6 | No |
| GLM-5.2 (max) | 91.2 | morph-glm53-744b |
| Claude Opus 4.8 (max) | 91.0 | No |
| Model | GPQA Diamond | Served on Morph |
|---|---|---|
| Gemini 3.1 Pro | 94.3 | No |
| Claude Opus 4.8 | 93.6 | No |
| GPT-5.5 | 93.6 | No |
| MiniMax M3 | 93 | morph-minimax3-428b |
| GLM-5.2 | 91.2 | morph-glm53-744b |
| DeepSeek-V4-Pro | 90.1 | No |
| Qwen3.7-Max | 90 | No |
| GLM-5.1 | 86.2 | No |
| Model | GPQA Diamond | Served on Morph |
|---|---|---|
| Qwen3.5-397B-A17B | 88.4 | morph-qwen35-397b |
| Qwen3.6-27B | 87.8 | morph-qwen36-27b |
| Claude 4.5 Opus | 87.0 | No |
| Qwen3.6-35B-A3B | 86.0 | No |
| Qwen3.5-27B | 85.5 | No |
| Gemma4-31B | 84.3 | morph-gemma4-31b |
The open-weight gap on GPQA Diamond is small. Kimi K3 and MiniMax M3 report within a point of GPT-5.5 and Claude Opus 4.8 in their respective cards, and a 27B dense model (Qwen3.6-27B) reports 87.8, above the 87.0 its card lists for Claude 4.5 Opus. The full benchmark tables for each model live on the Kimi K3, MiniMax M3, GLM-5.2, and Qwen 3.6 pages, with Anthropic models on Claude benchmarks.
Why GPQA Diamond Scores Disagree Between Reports
“95% confidence interval for a single GPQA Diamond run at 90% accuracy, from 198 questions.”
One question is 1/198 of the score, or about 0.5 points. At 90% accuracy the standard error of a single run is the square root of 0.9 × 0.1 / 198, about 2.1 points, so the 95% interval spans roughly 4.2 points either side. Two models 2 points apart on one run each cannot be ranked from that evidence alone.
Harness choices add a second source of spread. Claude Opus 4.8 appears at 93.6 in the GLM-5.2 card and at 91.0 (max effort) in the Kimi K3 card. GPT-5.5 appears at 93.6 and 93.5, which agree. The differences come from reasoning effort, temperature and top-p (Moonshot reports temperature 1.0 and top-p 0.95 for GPQA), prompt wording, and how the final letter is parsed out of a long reasoning trace. Epoch AI uses the simple-evals prompt, which asks the model to end with a line of the form "ANSWER: LETTER".
Compare models inside one model card or one independent leaderboard. Treat gaps under 4 points from single runs as ties. Check the reasoning-effort setting, since the same model at max and at default effort can differ by more than the gap to its nearest competitor.
Is the GPQA Benchmark Saturated?
For separating frontier models, mostly yes. The top independent score (96.3%) is above the 81.3% expert-validator accuracy on Diamond and well above the 69.7% that OpenAI's recruited PhD panel scored. The paper itself reports expert accuracy rising from 65% to 74% once experts' acknowledged mistakes are discounted, so some residual disagreement over answers is part of the dataset. The last few points sit inside both that disagreement and the 4-point noise band.
GPQA Diamond still works as a floor check. A model under 80% in 2026 is well behind the open-weight frontier on graduate-level science recall and reasoning, and a quick Diamond run catches quantization or serving regressions, since it is short, cheap, and has a fixed answer key. Contamination is the other caveat: the dataset ships with a canary string and a password-protected archive to keep it out of training crawls, but nothing guarantees every lab filtered it.
How to Run GPQA Diamond (lm-evaluation-harness)
The dataset is gated on Hugging Face at Idavidrein/gpqa; accept the terms there and log in with huggingface-cli login before running. EleutherAI's lm-evaluation-harness ships every split in zero-shot, n-shot, generative, and chain-of-thought variants (gpqa_diamond_zeroshot, gpqa_diamond_cot_zeroshot, gpqa_main_n_shot, and so on). For chat models served behind an OpenAI-compatible API, the chain-of-thought generative task is the closest match to how labs report the number.
pip install "lm_eval[api]"
huggingface-cli login
export OPENAI_API_KEY=YOUR_MORPH_API_KEY
lm_eval --model local-chat-completions \
--model_args model=morph-kimik3,base_url=https://api.morphllm.com/v1/chat/completions,num_concurrent=8,max_retries=3 \
--tasks gpqa_diamond_cot_zeroshot \
--apply_chat_template \
--log_samples \
--output_path results/Reasoning models write long traces before the answer letter. If the generation limit cuts the trace off, the item scores as wrong, so raise the limit before blaming the model. Keep the per-sample logs: a GPQA number without them cannot be checked for answer-parsing failures.
Cost is dominated by output tokens. At an assumed 10,000 output tokens per question, one Diamond pass on morph-kimik3 ($14.00/M output) costs about $27.72. Run it three to five times and average if you need to separate models closer than 4 points.
GPQA vs SWE-bench Pro, Terminal-Bench, and Other LLM Benchmarks
| Benchmark | What it measures | Format | Morph guide |
|---|---|---|---|
| GPQA Diamond | Graduate-level science (bio, physics, chem) | 198 four-option questions | This page |
| SWE-bench Pro | Resolving real repository issues | Agentic patch generation | SWE-bench Pro guide |
| Terminal-Bench | Completing tasks in a shell | Agentic, containerized | Claude benchmarks |
| lm-eval-harness tasks | MMLU-Pro, BBH, MATH, IFEval, GPQA | Academic multiple choice and generation | Eval harness guide |
GPQA tells you how much graduate science a model can recall and reason through in a single answer. It says little about coding agents. For that, read SWE-bench Pro and the coding sections of the Claude benchmarks page, and for picking an open model to self-host, the open-source LLM guide.
GPQA Benchmark FAQ
What is the GPQA benchmark?
GPQA (Graduate-Level Google-Proof Q&A) is a multiple-choice benchmark of 448 questions in biology, physics, and chemistry, written by domain experts and released by Rein et al. in November 2023 (arXiv 2311.12022). PhD-level experts score 65% on it (74% after discounting acknowledged mistakes), while skilled non-experts with unrestricted web access score 34%.
What is GPQA Diamond?
GPQA Diamond is the 198-question subset of GPQA where both expert validators answered correctly and the majority of non-expert validators answered incorrectly. The paper reports 81.3% expert accuracy and 22.1% non-expert accuracy on Diamond, though both are skewed by the selection rule. It is the split model cards and leaderboards report.
What is a good GPQA Diamond score?
Random guessing scores 25%. GPT-4 scored 39% on the main set in the original paper, and a panel of PhD experts recruited by OpenAI scored 69.7% on Diamond. Frontier models in 2026 report 91% to 94%, and the top model measured independently by Artificial Analysis scores 96.3%. A gap under about 4 points between two single runs is within statistical noise.
Which model has the highest GPQA Diamond score?
On the Artificial Analysis independent leaderboard (615 models evaluated), GPT-6 Astra at xhigh reasoning leads at 96.3%, followed by GPT-6 Astra at max reasoning at 96.1% and Gemini 3.8 Flash (High) at 95.3%. Among open-weight models, Kimi K3 reports 93.5 and MiniMax M3 is listed at 93 in the GLM-5.2 model card.
Why do different sources report different GPQA scores for the same model?
Prompt format, reasoning effort, sampling temperature, answer extraction, and the number of runs averaged all differ between labs. Claude Opus 4.8 is listed at 93.6 in the GLM-5.2 model card and at 91.0 in the Kimi K3 model card. With only 198 questions, that 2.6-point gap is about five questions.
How do I run GPQA Diamond myself?
Accept the terms on the gated Hugging Face dataset Idavidrein/gpqa, log in with huggingface-cli, and run EleutherAI's lm-evaluation-harness with a task such as gpqa_diamond_cot_zeroshot. Any OpenAI-compatible endpoint works through the local-chat-completions model type, including Morph's api.morphllm.com/v1.
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.
Run GPQA Diamond against open models on Morph
Kimi K3, MiniMax M3, GLM, and Qwen 3.6 behind one OpenAI-compatible endpoint. Point lm-evaluation-harness at api.morphllm.com/v1 and benchmark them on your own harness settings.