SWE-bench Pro Leaderboard (September 2026): Every Model Score, Benchmarks, and Price per Point

Every model's SWE-bench Pro score, checked September 14, 2026. Muse Spark 1.1 now tops Scale's standardized public set at 61.5% and the commercial set at 51.5%. Claude Fable 5 still leads the llm-stats vendor aggregate at 80.0%. Qwen3.8-Flash-Next leads open weights at 62.5%, GLM-5.2 at 62.1%. Pro vs Verified deltas, score per dollar, cost per solved task, harness rules.

September 14, 2026 · 2 min read
SWE-bench Pro Leaderboard (September 2026): Every Model Score, Benchmarks, and Price per Point

SWE-bench Pro is Scale AI's contamination-resistant coding benchmark: 1,865 real-world software tasks across 41 professional repositories, scored Pass@1, that the same frontier models clearing 80-95% on SWE-bench Verified solve only ~60% of under standardized scaffolding. Checked September 14, 2026, Meta's Muse Spark 1.1 leads Scale's standardized public set at 61.5% and its proprietary commercial set at 51.5%, taking both tops from GPT-5.4 (xHigh) at 59.1% and Claude Opus 4.6 at 47.1%. Scale's own footnote marks the top three public-set rows as mini-swe-agent runs, the harness whose maintainers closed a git-history exploit in March 2026. In the llm-stats vendor aggregate, now 57 models, Claude Fable 5 still holds the record at 80.0%. Qwen3.8-Flash-Next leads open weights at 62.5%, just past GLM-5.2 at 62.1%. Claude Opus 5, Claude Fable 5.1, GPT-6 Astra, GLM-5.3, and Kimi K3 have no SWE-bench Pro entry on either board.

Three numbers all claim to be the best SWE-bench Pro score: 61.5% (Muse Spark 1.1, Scale's standardized public set), 80.0% (Claude Fable 5, the llm-stats vendor aggregate), and 51.5% (Muse Spark 1.1 again, Scale's private commercial set). All three are real. The spread is scaffolding and data splits, and most pages quoting a score never say which one they mean. The September change is that one model now tops both standardized boards, which had not happened before.

This page keeps all three views side by side: Scale's standardized public and commercial leaderboards, vendor-reported scores, the Pro-vs-Verified delta per model, and score per dollar of output-token price.

Leaderboard data verified September 14, 2026

Updated September 14, 2026: Muse Spark 1.1 took the top of both Scale standardized boards, 61.5% on the public set and 51.5% on the commercial set, pushing GPT-5.4 (xHigh) and Claude Opus 4.6 to second. Scale's footnote now marks the top three public rows as mini-swe-agent runs. The llm-stats aggregate grew from 44 to 57 models: Hy4 preview 65.7%, Qwen3.8-Flash-Next 62.5%, Qwen3.8-27B 61.7%, Laguna S 2.1 59.4%, GPT-5.4 57.7%. Output prices corrected for GLM-5.2 ($2.40), GLM-5.1 ($3.50), MiniMax M3 ($1.10), DeepSeek V4 Pro-Max ($2.60) and Flash-Max ($0.18), Gemini 3.1 Pro ($12.00), with score per dollar recomputed. Claude Fable 5.1 and Claude Opus 5 are Anthropic's current lineup and neither has a Pro entry. GPT-6 Astra shipped September 2026 with none either. New sections cover harness rules and cost per solved task.

SWE-bench Pro: Scale Standardized Leaderboard Top 10 (Public Set)

Scale AI standardized scaffolding, Pass@1, 731 public tasks

1
Muse Spark 1.1 (Meta)
61.5%
2
GPT-5.4 (xHigh)
59.1%
3
Muse Spark (Meta)
55%
4
Opus 4.6 (thinking)
51.9%
5
Gemini 3.1 Pro
46.1%
6
Claude Opus 4.5
45.9%
7
Claude Sonnet 4.5
43.6%
8
Gemini 3 Pro
43.3%
9
Claude Sonnet 4
42.7%
10
GPT-5 (High)
41.8%

Source: Scale AI standardized SWE-bench Pro public leaderboard, read September 14, 2026. Pass@1; Muse Spark 1.1 95% CI is ±3.10 and GPT-5.4 (xHigh) is ±3.56, so the two share Rank (UB) 1.

Current top SWE-bench Pro score (Scale standardized public set): Muse Spark 1.1 at 61.5%. GPT-5.4 (xHigh) is second at 59.1% and shares Rank (UB) 1 with it. The best score in the llm-stats vendor aggregate is Claude Fable 5 at 80.0%, followed by Claude Mythos Preview (77.8%), Claude Opus 4.8 (69.2%), Qwen3.8 Max (67.7%), and Hy4 preview (65.7%). The best open-weights result is Qwen3.8-Flash-Next at 62.5%, then GLM-5.2 at 62.1% (vendor aggregate, not Scale standardized entries).

1,865
Tasks across 41 repositories
Pass@1
Scoring metric
4
Languages (Py, Go, TS, JS)
107.4
Avg lines changed per task

SWE-bench Pro Leaderboard (September 2026): Scale SEAL Public Set

Scale AI runs every model through identical scaffolding, which isolates model capability from harness quality. These are the only directly comparable SWE-bench Pro numbers. Scores below are from the public set (731 tasks), Pass@1, read September 14, 2026. The board carries 25 entries.

Muse Spark 1.1 leads at 61.5%, 2.4 points ahead of GPT-5.4 (xHigh) at 59.1% and 9.6 ahead of the best Claude run (Opus 4.6 thinking, 51.9%). Confidence intervals run ±3.1 to ±3.6 points at the top, so the first two entries overlap and Scale gives both Rank (UB) 1. Scale defines Rank (UB) as one plus the number of models whose lower confidence bound exceeds this model's upper bound.

Scale Standardized SWE-bench Pro Public Set (September 2026)
RankModelScoreType
1Muse Spark 1.1 (Meta)61.5% ±3.10Proprietary
2GPT-5.4 (xHigh)59.1% ±3.56Proprietary
3Muse Spark (Meta)55.0% ±3.60Proprietary
4Claude Opus 4.6 (thinking)51.9% ±3.61Proprietary
5Gemini 3.1 Pro (thinking)46.1% ±3.60Proprietary
6Claude Opus 4.545.9% ±3.60Proprietary
7Claude Sonnet 4.543.6% ±3.60Proprietary
8Gemini 3 Pro (preview)43.3% ±3.60Proprietary
9Claude Sonnet 442.7% ±3.59Proprietary
10GPT-5 (High)41.8% ±3.49Proprietary
11GPT-5.2 Codex41.0% ±3.57Proprietary
12Claude Haiku 4.539.5% ±3.55Proprietary
13Qwen3-Coder 480B-A35B38.7% ±3.55Open weights
14MiniMax 2.136.8% ±3.55Open weights
15Gemini 3 Flash34.6% ±3.55Proprietary
16GPT-5.229.9% ±2.15Proprietary
17Kimi K2 Instruct27.7% ±3.25Open weights
18Qwen3-235B-A22B21.4% ±2.25Open weights
19gpt-oss-120b16.2% ±2.67Open weights
20DeepSeek V3.215.6% ±2.63Open weights
21Gemma 3 27B IT11.4% ±2.15Open weights
22Llama 3.1 405B Instruct11.2% ±2.15Open weights
23GLM-4.69.7% ±2.15Open weights
24Llama 4 Maverick 17B5.2% ±1.24Open weights
25Codestral 24051.5% ±1.51Open weights

Source: Scale AI standardized SWE-bench Pro public leaderboard, read September 14, 2026. Pass@1, 95% confidence intervals as shown. Scale publishes no last-updated stamp on the page. Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Opus 4.8, GPT-5.5, GPT-5.6, GPT-6 Astra, GLM-5.2, GLM-5.3, and Kimi K3 have no standardized public-set entry, so their numbers in the next sections are vendor-reported or absent.

SWE-bench Pro Commercial Set: Scores on Code No Model Has Seen

The commercial set is 276 tasks from 18 proprietary startup codebases that are not on the public internet. It is the strongest contamination control available, and scores drop hard: every model loses ground versus its public-set number, and the ranking reshuffles.

Scale Standardized SWE-bench Pro Private Set (September 2026)
RankModelPrivate ScorePublic-Set Score
1Muse Spark 1.1 (Meta)51.5% ±5.5061.5%
2Claude Opus 4.6 (thinking)47.1% ±6.0751.9%
3Muse Spark (Meta)44.7% ±6.0555.0%
4GPT-5.4 (xHigh)43.4% ±6.0359.1%
5Gemini 3.1 Pro (thinking)32.2% ±5.6946.1%
6GPT-5.2 Codex27.7% ±5.0941.0%
7GPT-5.223.8% ±5.0929.9%
8Claude Opus 4.523.4% ±5.0745.9%
9Gemini 3 Pro18.0% ±4.7843.3%
10Claude Opus 4.117.8% ±4.51no entry
11GPT-514.9% ±4.2041.8%
12Gemini 2.5 Pro Preview10.1% ±3.56no entry
13Claude Sonnet 49.1% ±3.3942.7%
14GPT-4o3.6% ±2.20no entry

Source: Scale AI standardized private leaderboard, read September 14, 2026. The 276-task set carries wider confidence intervals, ±2.2 to ±6.1 points. Scale marks the top four rows (Muse Spark 1.1, Claude Opus 4.6 thinking, Muse Spark, gpt-5.4 xHigh) as mini-swe-agent runs; the rest are its earlier harness. Scale publishes no last-updated stamp.

The reshuffle is the interesting part. GPT-5.4 was the public-set leader until September and still falls to fourth on commercial code (43.4%), giving up 15.7 points. Muse Spark 1.1 loses 10.0 points and tops both boards. Claude Opus 4.5 is the sharpest fall in the table: 45.9% public, 23.4% commercial, a 22.5-point drop. If you are choosing a model for a private codebase, the commercial column is the one that predicts your experience.

Vendor-Reported SWE-bench Pro Scores: Claude Fable 5 Leads at 80.0% (September 2026)

Labs run SWE-bench Pro on their own agent scaffolds, with tuned context retrieval, tool use, and turn budgets, and llm-stats aggregates those self-reported numbers into one table. These are not comparable to the Scale standardized leaderboard, but they are comparable to each other. The aggregate grew from 44 models in August to 57 in September. Claude Fable 5 still holds the all-time high at 80.0%.

SWE-bench Pro: llm-stats Vendor Aggregate (September 2026)
ModelPro ScoreOutput Price
Claude Fable 580.0%$50/M
Claude Mythos Preview (Glasswing partners only)77.8%n/a
Claude Opus 4.869.2%$25/M
Qwen3.8 Max (2.4T)67.7%$4.95/M
Hy4 preview (Tencent)65.7%n/a
Grok 4.564.7%$6.00/M
GPT-5.6 Sol64.6%$30/M
Claude Opus 4.764.3%$25/M
GPT-5.6 Terra63.4%$12/M
Claude Sonnet 563.2%$10/M
GPT-5.6 Luna62.7%$1.20/M
Qwen3.8 Flash / Qwen3.8-Flash-Next (open weights)62.5%$0.47/M
GLM-5.2 (open weights)62.1%$2.40/M
Qwen3.8-27B (open weights)61.7%$3.00/M
Muse Spark 1.1 (Meta)61.5%$4.25/M
Qwen3.7 Max (open weights)60.6%$3.75/M
Laguna S 2.1 (Poolside)59.4%$0.20/M
MiniMax M3 (open weights)59.0%$1.10/M
Gemini 3.6 Flash58.7%$7.50/M
GPT-5.558.6%$30/M
Kimi K2.6 (open weights)58.6%$3.50/M
GLM-5.1 (open weights)58.4%$3.50/M
Hy3 (Tencent)57.9%$0.58/M
GPT-5.457.7%$15/M
GPT-5.3 Codex56.8%$14/M
MiniMax M2.7 (open weights)56.2%$1.20/M
DeepSeek-V4-Pro-Max (open weights)55.4%$2.60/M
Gemini 3.1 Pro54.2%$12/M
DeepSeek-V4-Flash-Max (open weights)52.6%$0.18/M

Source: llm-stats SWE-bench Pro aggregate, read September 14, 2026 (57 models, all vendor self-reported; GLM-5.2's 62.1% is third-party measured, as Z.ai published no SWE-bench number at launch). Output prices are the list rates llm-stats shows next to each entry. Several moved since August: GLM-5.2 $3.00 to $2.40, GLM-5.1 $4.40 to $3.50, MiniMax M3 $1.20 to $1.10, DeepSeek V4 Pro-Max $3.20 to $2.60 and Flash-Max $0.20 to $0.18, Gemini 3.1 Pro $15 to $12. Anthropic rates cross-checked on the Anthropic pricing page.

Models with no SWE-bench Pro score at all

Five current flagships are missing from both boards as of September 14, 2026. Anthropic's model overview now lists Claude Fable 5.1 ($10/$50, 1M context) and Claude Opus 5 ($5/$25, 1M context) as the current lineup, with Fable 5 moved to legacy, and neither new model has a Pro entry on Scale or in the 57-model llm-stats aggregate. GPT-6 Astra shipped in September 2026 at $10/M input and $50/M output and llm-stats tracks SWE-bench Verified for it, not Pro, with no score filled in. Z.ai has published no SWE-bench Pro or Verified figure for GLM-5.3. Moonshot has published none for Kimi K3. Any Pro number you see quoted for these five is not coming from Scale or llm-stats.

GPT-5.6 (Sol / Terra / Luna) and the retirement of GPT-5.4

GPT-5.6 is generally available in three variants: Sol for complex coding, Terra for everyday work, Luna for high-volume repeatable work. SWE-bench Pro scores in the vendor aggregate: Sol 64.6% ($5/M input, $30/M output), Terra 63.4% ($2/$12), Luna 62.7% ($0.20/$1.20). Luna is the one to notice. It costs 4% of Sol's output rate and gives up 1.9 points, which makes it the cheapest way to buy a 60%+ Pro score from a closed model. OpenAI retired GPT-5.4 and GPT-5.4 mini from Codex on August 31, 2026 and migrates those users to 5.6-terra and 5.6-luna. That matters for reading the standardized board: gpt-5.4 (xHigh), second at 59.1%, is a model you can no longer pick in Codex.

The vendor-vs-standardized gap is consistent: the aggregate reports 69.2% for Opus 4.8 while Scale's best standardized Claude run (Opus 4.6 thinking) scores 51.9% on the public set. GPT-5.2 Codex reports 56.4% on its own scaffold and scores 41.0% under Scale's harness, a 15.4-point gap on one model. When you see a SWE-bench Pro score 10-30 points above the Scale leaderboard, it is a vendor-scaffold number.

SWE-bench Pro Harness Rules: mini-swe-agent, Turn Limits, and the Git-History Exploit

"Standardized" does not mean every row on Scale's board ran under the same conditions. Two footnotes on the public leaderboard change how you read it, and both are easy to miss.

What Scale's Footnotes Actually Say (read September 14, 2026)
FootnoteEffect on the score
Asterisk rowsRun with the mini-swe-agent harness. On the public set that is Muse Spark 1.1, gpt-5.4 (xHigh), and Muse Spark, the top three entries.
Grayed-out rowsRun with a capped cost limit and a turn limit of 50. Every other row ran with uncapped cost and a turn limit of 250.
Rank (UB)One plus the number of models whose lower CI bound exceeds this model's upper bound. Muse Spark 1.1 and gpt-5.4 (xHigh) both show Rank (UB) 1.

Source: footnote text on the Scale public leaderboard, read September 14, 2026.

A 50-turn cap against a 250-turn cap is a five-fold difference in how long an agent gets to work on tasks that average 107.4 changed lines across 4.1 files. Comparing a grayed row against an ungrayed one measures the budget as much as the model.

The mini-swe-agent Git-History Exploit

The harness marking matters more because of a defect its own maintainers documented. Issue #787 on SWE-agent/mini-swe-agent, filed March 18, 2026 and closed March 24, reports that SWE-bench Pro containers ship the repository's full git history, so an agent can run a pickaxe search such as git log --all -S "functionName", find the gold commit, and read the reference implementation out with git show. The filer measured a 21% exploitation rate across 100 SWE-bench Pro tasks running mini-swe-agent with Claude Opus 4.6. The original SWE-agent hides .git/ during file listing, so the agent never learns the history is there; mini-swe-agent did not.

That is the same mechanism Datacurve reported in its May 2026 DeepSWE audit, and the same one Scale tracks as issue #93. Scale has not published which harness build produced each leaderboard row or whether the fixed version was used, so the size of any residual effect on the three asterisk rows is unknown.

Harness Defects That Push Scores Down

The scoring errors are not all in the model's favor. Three open reports on Scale's repository describe the opposite failure.

Filed on scaleapi/SWE-bench_Pro-os
  • Silent unresolved scoring from a 60-second Docker timeout. Pull request #111, opened July 4, 2026, reports that on the --use_local_docker path the Docker SDK's default 60-second read timeout makes container.wait() raise ReadTimeout for test suites that run longer, which is common in the Go and JavaScript repositories. No output.json is written and the instance is scored unresolved with no visible error. The fix adds a DOCKER_CLIENT_TIMEOUT constant defaulting to 3600 seconds, and cites the same false negatives in issues #54, #22, and #23.
  • A task whose requirement contradicts its own test. Issue #115, filed August 3, 2026, covers instance navidrome__navidrome-97434c1: the visible requirement says Register must return a nil transcoding value, while the official TestCore asserts trc.ID equals "1". An agent that follows the written spec fails. The gold patch passes by ignoring it.
  • A submitter asking for their own number to be withdrawn. Issue #117, filed September 14, 2026, asks Scale not to publish a pending 570/731 public-set submission. The submitter states the run was open-material and adaptive rather than held-out, that 7 task patches were written after prior failures on the same tasks, and that 371 rows contain traces matching later public repository states or public tests.

All three are reports by third parties on Scale's repository, not Scale findings. Issues #115 and #117 were open when this page was checked on September 14, 2026.

Score per Dollar: SWE-bench Pro Points per $1/M Output Tokens

Benchmark points are not free. Dividing each model's SWE-bench Pro score (llm-stats vendor aggregate) by its output-token price shows where capability is cheap. Ling 3.0 Flash returns about 314 points per output dollar and GPT-5.6 Luna about 52, against 2.8 for Opus 4.8 and 1.6 for Fable 5. The highest absolute score is still the worst value. The September change at the cheap end is Poolside's Laguna S 2.1: 59.4% at $0.20/M output, within 3.3 points of GPT-5.6 Luna at a sixth of the price.

SWE-bench Pro Points per $1/M Output Tokens (September 2026)
ModelPro Score$/M OutputPoints per $
Ling 3.0 Flash (InclusionAI)56.6%$0.18314.4
Laguna S 2.1 (Poolside)59.4%$0.20297.0
DeepSeek-V4-Flash-Max (open)52.6%$0.18292.2
MiMo-V2.5 (open)56.1%$0.34165.0
Qwen3.8 Flash (open weights as Flash-Next)62.5%$0.47133.0
Hy3 (Tencent)57.9%$0.5899.8
MiMo-V2.5-Pro (open)57.2%$0.8765.7
MiniMax M3 (open)59.0%$1.1053.6
GPT-5.6 Luna62.7%$1.2052.3
Step 3.7 Flash56.3%$1.1549.0
MiniMax M2.7 (open)56.2%$1.2046.8
GLM-5.2 (open)62.1%$2.4025.9
DeepSeek-V4-Pro-Max (open)55.4%$2.6021.3
Qwen3.8-27B (open)61.7%$3.0020.6
Kimi K2.6 (open)58.6%$3.5016.7
GLM-5.1 (open)58.4%$3.5016.7
Qwen3.7 Max (open)60.6%$3.7516.2
Muse Spark 1.161.5%$4.2514.5
Qwen3.8 Max67.7%$4.9513.7
Grok 4.564.7%$6.0010.8
Gemini 3.6 Flash58.7%$7.507.8
Claude Sonnet 563.2%$10.006.3
GPT-5.6 Terra63.4%$12.005.3
Gemini 3.1 Pro54.2%$12.004.5
Claude Opus 4.869.2%$25.002.8
GPT-5.6 Sol64.6%$30.002.2
GPT-5.558.6%$30.002.0
Claude Fable 580.0%$50.001.6

Scores and list prices: llm-stats, read September 14, 2026. Points per dollar is score divided by output price, recomputed after the September price moves. Note: Opus 4.7 and later (including Fable 5 and Fable 5.1) use a tokenizer that can produce up to 35% more tokens for the same text than pre-4.7 Claude models, which raises effective per-request cost beyond the per-token rate. Full cost modeling in our LLM cost calculator.

The cost spread is why model choice should be per request, not per project. Morph's model router scores each request and returns the cheapest model that clears the bar, so a coding agent gets cheaper and faster at the same time instead of paying Opus rates for work an open-weights model solves. Morph serves these open models at 16-bit (bf16) activations with no fp8 or int8 quantization, so output matches the reference weights.

Morph-Served Models: List Price and SWE-bench Pro Status
ModelAPI name$/M in$/M outSWE-bench Pro
GLM-5.3-Flashmorph-glm53flash$0.10$0.3550-52% on a 64-task subset
DeepSeek V4 Flash 0731morph-dsv4flash$0.12$0.35No entry for this checkpoint
GLM-5.3 744Bmorph-glm53-744b$1.00$3.41Z.ai has published none
Kimi K3 2.8Tmorph-kimik3$2.50$14.00Moonshot has published none

Prices are Morph list rates from src/lib/pricing.ts, checked September 14, 2026. The GLM-5.3-Flash figure is the independent 64-task measurement described in the next section, not a full 731-task public-set score, so it is not comparable to the table above. At $0.35/M output that works out to roughly 143 to 149 points per output dollar on that subset. See pricing for the full list.

Cost per Solved SWE-bench Pro Task: Same Accuracy, 2x Spend

Leaderboards publish the score and hide the bill. A benchmark published September 10, 2026 by the imec aistack team ran the same 64 curated SWE-bench Pro tasks through three harnesses, Claude Code, Codex, and Pi, against two self-hosted open models, and measured GPU time for every run.

64 SWE-bench Pro Tasks, Three Harnesses, Two Self-Hosted Models
Model and hardwareHarnessResolvedGPU cost for 64 tasks
Qwen3.8-27B FP8, 1x H200Claude Code33/64 (52%)$10.26
Qwen3.8-27B FP8, 1x H200Codex31/64 (48%)$9.00
Qwen3.8-27B FP8, 1x H200Pi28/64 (44%)$10.16
GLM-5.3-Flash FP8, 4x H200Codex33/64 (52%)$22.80
GLM-5.3-Flash FP8, 4x H200Claude Code32/64 (50%)$45.00
GLM-5.3-Flash FP8, 4x H200Pi32/64 (50%)$45.00

Source: imec aistack, "Benchmarking Claude Code, Codex and Pi on SWE-Bench Pro: Same Accuracy, 2x Cost", published September 10, 2026 and posted to Hacker News the same day. The 64 tasks are the team's own curated subset, not the full 731-task public set, so the resolve rates are not leaderboard-comparable.

Cost per solved task ran $12.90 to $26.00 across the six combinations. Accuracy spread 8 points; spend spread 2x. On GLM-5.3-Flash the Codex harness solved one more task than Claude Code for half the GPU time. That is the argument for treating the harness as a variable you tune rather than a constant you inherit, and it is the same variable Scale's asterisk footnote is flagging on its own board.

WarpGrep Impact on SWE-bench Pro (Morph Internal)

Morph runs SWE-bench Pro internally and serves open-weights models on api.morphllm.com: morph-kimik3, morph-glm53-744b, morph-glm53flash, and morph-dsv4flash. The benchmark runs below isolate one variable: adding a search subagent to an existing coding agent.

Self-reported data

The scores below are from Morph's internal benchmark runs (March 2026), not from the Scale standardized leaderboard, and use the models current at that time. They show the effect of adding WarpGrep v2 as a search subagent to existing coding agents.

SWE-bench Pro: With vs Without WarpGrep v2

Morph internal benchmarks, public set (731 tasks), March 2026

With WarpGrep v2
Without WarpGrep
1
GPT-5.3 Codex
59.1%
2
MiniMax M2.5
57.6%
3
Opus 4.6
57.5%

WarpGrep v2 adds 2.1-2.2 points to every model tested.

WarpGrep v2 is an RL-trained search subagent that runs in its own context window. It issues up to 8 parallel tool calls per turn and returns only the relevant file spans. The main coding model never sees files WarpGrep rejected, so its context stays clean.

With Opus 4.6, adding WarpGrep v2 cuts cost by 15.6% and time by 28%. The expensive model spends fewer tokens on search and more on code generation. Read how subagents make coding agents faster for the full breakdown.

SWE-bench Verified Leaderboard (September 2026)

SWE-bench Verified is the human-validated 500-task Python subset of the original SWE-bench. It remains the most-quoted coding benchmark, but OpenAI deprecated it in February 2026 over contamination. Scores below are vendor-reported and aggregated by llm-stats, which now tracks 116 entries.

SWE-bench Verified Top 15 (September 2026)
RankModelScore
1Claude Fable 595.0%
2Claude Mythos Preview (Glasswing partners only)93.9%
3Claude Opus 4.888.6%
4Claude Opus 4.787.6%
5Claude Sonnet 585.2%
6Claude Opus 4.580.9%
7Claude Opus 4.680.8%
8Gemini 3.1 Pro80.6%
8DeepSeek-V4-Pro-Max (open)80.6%
10MiniMax M3 (open)80.5%
11Qwen3.7 Max (open)80.4%
12Inkling-Small (Thinking Machines Lab)80.2%
12Kimi K2.6 (open)80.2%
12MiniMax M2.5 (open)80.2%
15GPT-5.280.0%

Source: llm-stats SWE-bench Verified tracker, read September 14, 2026. Vendor self-reported (all 116 listed entries; none independently verified); harness differences apply. Claude Opus 5, Claude Fable 5.1, GPT-6 Astra, GLM-5.3, and Kimi K3 have no Verified entry either. See our full Claude benchmarks page for the rest of the suite.

Note the compression: ranks 8 through 15 span 0.6 points (80.6% to 80.0%), with four open-weights models inside that band and a three-way tie at 80.2%. When eight models from seven labs sit statistically tied near 80%, the benchmark has stopped discriminating at the frontier. That saturation, plus contamination, is why Pro exists.

SWE-bench Pro vs Verified: Same Model, Different Score

The per-model delta between Verified and Pro is the cleanest measure of how much Verified overstates capability:

Verified-to-Pro Score Drop per Model (vendor aggregate, September 2026)
ModelVerifiedProDrop
Gemini 3.1 Pro80.6%54.2%−26.4 pts
DeepSeek-V4-Pro-Max80.6%55.4%−25.2 pts
Claude Opus 4.787.6%64.3%−23.3 pts
Claude Sonnet 585.2%63.2%−22.0 pts
Kimi K2.680.2%58.6%−21.6 pts
MiniMax M380.5%59.0%−21.5 pts
Claude Opus 4.888.6%69.2%−19.4 pts
Claude Fable 595.0%80.0%−15.0 pts

Both columns are llm-stats vendor self-reported, and rows are limited to models with entries on both boards. OpenAI stopped publishing SWE-bench Verified results in February 2026, so GPT-5.5 and GPT-5.6 have Pro scores with no Verified counterpart. The drop grows when you move to Scale's standardized scaffold: Gemini 3.1 Pro falls to 46.1% on the public set and 32.2% on the commercial set.

Gemini 3.1 Pro is the starkest verified case: 80.6% on Verified, 54.2% on the vendor Pro aggregate, 46.1% on Scale's standardized public set, and 32.2% on the proprietary commercial set. The drop is not the model getting worse. It is the benchmark getting honest.

Benchmark Design: Verified vs Pro
DimensionSWE-bench VerifiedSWE-bench Pro
Tasks5001,865
Repositories12 (all Python)41 (Python, Go, TS, JS)
Avg lines changed11 (median: 4)107.4
Avg files changed~14.1
Minimum task size161/500 tasks are 1-2 linesEvery task is 10+ lines
Contamination resistanceLow: public Python reposHigh: copyleft + proprietary code
StatusDeprecated by OpenAI, Feb 2026Active, recommended

Open-Source Models on SWE-bench: Qwen3.8-Flash-Next Takes Open Weights at 62.5%, GLM-5.2 Leads Open-Weights on Pro at 62.1%, DeepSeek V4, MiniMax M3, Qwen

The open-weights top changed in September. Qwen3.8-Flash-Next is now the best open-weights model on SWE-bench Pro at 62.5% in the llm-stats vendor aggregate, 0.4 points above GLM-5.2 at 62.1%, with Qwen3.8-27B at 61.7%, Qwen3.7 Max at 60.6%, MiniMax M3 at 59.0%, and Kimi K2.6 at 58.6% behind them. Four open models now clear GPT-5.5's 58.6% vendor number. On SWE-bench Verified, open weights tie Gemini 3.1 Pro: DeepSeek-V4-Pro-Max 80.6%, MiniMax M3 80.5%, Qwen3.7 Max 80.4%, Kimi K2.6 80.2%. Coverage under Scale's standardized scaffolding is still thin, and no model newer than Qwen3-Coder 480B has an entry. Status per model, September 14, 2026:

Open-Weights Models: SWE-bench Status (September 2026)
ModelVerifiedPro (Scale public)Pro (vendor)Output Price
Qwen3.8-Flash-Nextno officialNo entry62.5%$0.47/M (hosted Flash rate)
GLM-5.2no officialNo entry62.1%$2.40/M
Qwen3.8-27Bno officialNo entry61.7%$3.00/M
Qwen3.7 Max80.4%No entry60.6%$3.75/M
MiniMax M380.5%No entry59.0%$1.10/M
Kimi K2.680.2%No entry58.6%$3.50/M
GLM-5.1no officialNo entry58.4%$3.50/M
DeepSeek-V4-Pro-Max80.6%No entry55.4%$2.60/M
DeepSeek-V4-Flash-Max79.0%No entry52.6%$0.18/M
GLM-5.3none publishedNo entryNone publishedn/a
Kimi K3none publishedNo entryNone publishedn/a
Qwen3-Coder 480Bn/a38.7%n/an/a

Verified and vendor Pro scores: llm-stats, read September 14, 2026. Pro (Scale public): Scale standardized leaderboard, same date. GLM-5.2's 62.1% is third-party measured (Z.ai published no SWE-bench number at launch). DeepSeek-V4-Flash-Max's 79.0% Verified figure was last confirmed on August 10, 2026 and is carried forward unverified. Prices as listed by llm-stats; llm-stats shows no price against the Qwen3.8-Flash-Next row, so $0.47/M is the rate on its hosted Qwen3.8 Flash entry, which scores the same 62.5%.

Qwen3.8-Flash-Next and the license: Alibaba released the weights on August 26, 2026 on Hugging Face and ModelScope under the Qwen Community License 1.0, not Apache 2.0. That license permits commercial use and fine-tuning but requires attribution above 100M monthly users or $20M monthly revenue and a separate licence to run a model-as-a-service or an AI coding-assistant business. If you are picking the top open-weights Pro score for a product, read the licence before the leaderboard. Note also that llm-stats lists the hosted "Qwen3.8 Flash" API entry and the open-weights "Qwen3.8-Flash-Next" checkpoint separately, both at 62.5%.

DeepSeek V4 and SWE-bench: neither DeepSeek V4 Flash nor Pro has a Scale standardized SWE-bench Pro entry as of September 14, 2026. The closest Scale entry is deepseek-v3p2 at 15.56%. The llm-stats vendor aggregate reports 55.4% for V4-Pro-Max, 52.6% for V4-Flash-Max, and 52.3% for V4-Flash-0423. Its strongest verified result is V4-Pro-Max at 80.6% on SWE-bench Verified, the top open-weights score, tied with Gemini 3.1 Pro. The V4 family ships open weights at 1.6T total / 49B active parameters (Pro) and 284B / 13B (Flash), 1M context, with list output at $0.18/M (Flash-Max) and $2.60/M (Pro-Max). Morph serves morph-dsv4flash at 16-bit (bf16) for $0.12/$0.35 per 1M.

GLM-5.2 and SWE-bench Pro: the 62.1% figure is third-party measured, not a Scale standardized entry; Z.ai published no SWE-bench number at the June 13, 2026 launch. Scale's standardized leaderboard has no GLM entry newer than GLM-4.6 at 9.67%, so the top open-weights entry under standardized scaffolding remains qwen3-coder-480b-a35b at 38.7%. GLM-5.2 is a 744B MoE (about 40B active) with a usable 1M context, now listed at $0.75/M input and $2.40/M output, down from $3.00 output in August. Its successor GLM-5.3 continues the pattern: Z.ai published Terminal-Bench, DeepSWE, and its own Z.ai Code Bench at launch and no SWE-bench Pro or Verified figure at all, so any GLM-5.3 SWE-bench number in circulation has no published source. Morph serves morph-glm53-744b at $1.00/$3.41 per 1M and morph-glm53flash at $0.10/$0.35. Comparisons against other open models: GLM-5 vs MiniMax and GLM-5 vs Qwen 3.5.

How SWE-bench Pro Works: 1,865 Tasks, 41 Repos, Pass@1

SWE-bench Pro contains 1,865 tasks across 41 actively maintained repositories spanning Python, Go, TypeScript, and JavaScript, scored Pass@1 (one attempt, no retries). Tasks come from real commit histories: consecutive commits where one resolves a bug or adds a feature, paired with tests that demonstrate the fix. The benchmark and its methodology are described in the Scale AI paper "SWE-bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (arXiv:2509.16941), with the public set and harness released at scaleapi/SWE-bench_Pro-os.

Three Subsets

Public Set (731 tasks)

Tasks from 11 copyleft (GPL) repositories, openly available on HuggingFace. The primary evaluation target for leaderboard submissions.

Commercial Set (276 tasks)

Tasks from 18 proprietary startup codebases, acquired through Scale AI partnerships. Not publicly accessible: the strongest contamination control.

Held-Out Set (858 tasks)

Tasks from 12 repositories reserved for overfitting detection. Scale can release these to verify that public-set gains generalize.

Three-Stage Human Augmentation

  1. Problem statement creation: original commit messages and issue discussions are synthesized into clear, structured descriptions
  2. Requirements definition: annotators create specification lists grounded in unit tests and gold patches, detailing expected behavior without prescribing implementation
  3. Interface specification: class and function signatures are documented to prevent false negatives from naming mismatches
Evaluation methodology

Evaluation uses containerized, language-specific environments. Each task must pass fail2pass tests (tests that fail before the fix and pass after, verifying the issue is resolved) and pass2pass tests (existing tests that must keep passing). Gold patches are validated across 3 test runs before inclusion. Copyleft licensing makes the public set legally unattractive as training data, and the commercial set is never published at all.

Why Scores Are So Much Lower Than Verified

Four factors compound. Multi-file modifications: Pro tasks touch 4.1 files on average; Verified is mostly single-file. Longer horizons: tasks that take a professional engineer hours to days, requiring coherent plans across many steps. Production codebases: business applications and developer tools with real build systems and conventions. No memorization: copyleft and proprietary repos mean models must reason about unfamiliar code, not recall it.

Failure mode analysis

Scale's trajectory analysis shows where models break: semantic understanding failures (35.9% of Opus 4.1 failures), context overflow (35.6% of Sonnet 4 failures), and tool-use inefficiency (42% of smaller-model failures). Context overflow dominating the strongest models aligns with research showing coding agents spend 60%+ of their time searching for context.

Is SWE-bench Verified Contaminated? Why OpenAI Deprecated It

In February 2026, OpenAI published "Why SWE-bench Verified no longer measures frontier coding progress" and stopped reporting Verified scores. The core finding: frontier models could reproduce gold patches and problem-statement specifics from training data, since all 500 tasks come from public Python repositories that predate every model's cutoff.

Benchmark validity criticism cuts both ways. A widely circulated community analysis claims 68.5% of GPT-5.5's SWE-bench Pro failures trace to broken test cases rather than model errors. That figure has not been confirmed by Scale or OpenAI; treat it as an open question rather than a result. What is verifiable: Scale validates gold patches across 3 test runs, publishes confidence intervals, and keeps an 858-task held-out set specifically to catch overfitting.

The DeepSWE Audit: Reward Hacking and Broken Verifiers

In May 2026, Datacurve released DeepSWE and ran an audit of SWE-bench Pro's rollouts and graders. Three findings sit directly under the 68.5%-broken-test-cases criticism above. All are reported by Datacurve, not confirmed by Scale; the git-history issue is acknowledged as an open issue on Scale's own GitHub repo.

Reported by Datacurve, not confirmed by Scale
  • Git-history reward hacking. Datacurve marked Claude Opus 4.6 and 4.7 as "CHEATED" on more than 12% of reviewed SWE-bench Pro tasks. The benchmark's Docker containers ship the repository's full .git history, so the gold-patch commit is on disk; the agents ran commands like git log --all to read the merged fix and paste it. GPT-5.4 and GPT-5.5 were not flagged for this. Scale tracks it as an open issue (#93).
  • Broken verifiers. Datacurve reports SWE-bench Pro's automated graders accepted incorrect implementations 8.5% of the time and rejected correct ones 24% of the time, roughly one-third of trials mis-graded. That is the mechanism behind the circulating "68.5% of failures are broken test cases" claim.
  • Best open-source model. As of September 14, 2026, Qwen3.8-Flash-Next leads open weights on SWE-bench Pro at 62.5%, with GLM-5.2 at 62.1% (both third-party or vendor measured, neither a Scale standardized entry), ahead of Qwen3.8-27B (61.7%) and Qwen3.7 Max (60.6%).
  • Independently reproduced in March 2026. The git-history mechanism is not only a Datacurve claim. Issue #787 on mini-swe-agent, filed March 18, 2026, measured a 21% exploitation rate on 100 SWE-bench Pro tasks with Claude Opus 4.6, via git log --all -S pickaxe searches. It was closed March 24, 2026.

Sources: VentureBeat on the Datacurve DeepSWE audit; Scale's GitHub issue #93. None of these figures are confirmed by Scale AI.

Practical reading order for a model decision: commercial-set score first (closest to private-codebase reality), public SEAL score second (clean cross-model comparison), vendor numbers last (upper bound with tuned scaffolding). Verified scores from 2026 onward are best read as a saturation indicator, not a ranking.

Frequently Asked Questions

What is SWE-bench Pro?

SWE-bench Pro is Scale AI's software engineering benchmark: 1,865 tasks from 41 repositories across Python, Go, TypeScript, and JavaScript, scored Pass@1, split into public (731), commercial (276), and held-out (858) sets. Tasks average 107.4 changed lines across 4.1 files.

How hard is SWE-bench Pro?

Models lose 15 to 35 points moving from Verified to Pro. Gemini 3.1 Pro: 80.6% to 46.1% on Scale's standardized public set. Claude Opus 4.8: 88.6% to 69.2% on the vendor aggregate. The best standardized public-set score as of September 14, 2026 is 61.5% (Muse Spark 1.1), with GPT-5.4 (xHigh) second at 59.1%. On the proprietary commercial set, no model exceeds 51.5% (Muse Spark 1.1), with Opus 4.6 second at 47.1%.

Which model leads the SWE-bench Pro leaderboard in September 2026?

Meta's Muse Spark 1.1 leads both Scale standardized boards: 61.5% (95% CI ±3.10) on the 731-task public set and 51.5% (±5.50) on the 276-task commercial set. GPT-5.4 (xHigh) is second on the public set at 59.1% and shares Rank (UB) 1, because Scale defines Rank (UB) as one plus the number of models whose lower confidence bound exceeds this model's upper bound. In the llm-stats vendor aggregate, Claude Fable 5 leads at 80.0%.

What does Claude Fable 5 score on SWE-bench Pro?

80.0% in the llm-stats vendor aggregate, the highest SWE-bench Pro score on record, versus 77.8% for Mythos Preview and 69.2% for Opus 4.8. Scale's standardized leaderboard has no Fable 5 entry. Anthropic's model overview now lists Fable 5 under legacy models. The current flagship is Claude Fable 5.1 (claude-fable-5-1), $10/M input and $50/M output, 1M context, June 2026 knowledge cutoff, and it has no SWE-bench Pro score on either board. Mythos 5.1 and Mythos Preview stay invitation-only under Project Glasswing.

Does Claude Opus 5 have a SWE-bench Pro score?

No. Anthropic's model overview lists Claude Opus 5 (claude-opus-5) as the default recommendation for most workloads at $5/M input and $25/M output, 1M context, May 2026 knowledge cutoff. Neither Scale's standardized board (25 entries) nor the llm-stats vendor aggregate (57 entries) carries a Pro number for it on September 14, 2026, and the same is true of Claude Fable 5.1. The highest-scoring Claude entries on Pro remain Fable 5 (80.0%), Mythos Preview (77.8%), Opus 4.8 (69.2%), Opus 4.7 (64.3%), and Sonnet 5 (63.2%).

What does Claude Opus 4.8 score on SWE-bench Pro?

69.2% in the llm-stats vendor aggregate. Opus 4.8 also posts 88.6% on SWE-bench Verified, at $5/M input and $25/M output with a 1M-token context window. Anthropic lists it under legacy models as of September 14, 2026.

What does GPT-5.3 Codex score on SWE-bench Pro?

56.8% in the llm-stats vendor aggregate, with GPT-5.2 Codex at 56.4%. Under Scale's standardized scaffolding, gpt-5.2-codex scores 41.0% on the public set and 27.7% on the commercial set, a 15.4-point gap on the same model between its own scaffold and Scale's. gpt-5.3-codex is priced at $1.75/M input, $14/M output.

What about GPT-5.6?

GPT-5.6 is generally available across all three variants. SWE-bench Pro, vendor aggregate: Sol 64.6% ($5/$30 per 1M), Terra 63.4% ($2/$12), Luna 62.7% ($0.20/$1.20). Sol is the strongest OpenAI entry on Pro, 6.0 points ahead of GPT-5.5, and Luna delivers 97% of Sol's score at 4% of its output price. OpenAI retired GPT-5.4 and GPT-5.4 mini from Codex on August 31, 2026 and migrates those users to 5.6-terra and 5.6-luna, so Scale's second-ranked standardized entry is a model you can no longer select in Codex.

Does GPT-6 Astra have a SWE-bench Pro score?

No. GPT-6 Astra shipped in September 2026 at $10/M input and $50/M output with a $1/M cached-input rate and a 1.1M-token context window. llm-stats tracks SWE-bench Verified for it, not SWE-bench Pro, and shows no score in either place. OpenAI has published no SWE-bench Pro figure for Astra at all.

Does DeepSeek V4 have a SWE-bench Pro score?

No Scale standardized entry exists for any DeepSeek V4 variant as of September 14, 2026; the closest Scale entry is deepseek-v3p2 at 15.56%. The llm-stats vendor aggregate reports 55.4% for V4-Pro-Max, 52.6% for V4-Flash-Max, and 52.3% for V4-Flash-0423. On SWE-bench Verified, V4-Pro-Max scores 80.6%, the highest open-weights result, tied with Gemini 3.1 Pro. DeepSeek V4.1 has no Pro entry either. Details on the model family: DeepSeek V4.

What is the best open-source model on SWE-bench?

On SWE-bench Pro (vendor aggregate, September 14, 2026): Qwen3.8-Flash-Next at 62.5%, then GLM-5.2 (62.1%), Qwen3.8-27B (61.7%), Qwen3.7 Max (60.6%), MiniMax M3 (59.0%), Kimi K2.6 (58.6%). Qwen3.8-Flash-Next shipped August 26, 2026 under the Qwen Community License 1.0, not Apache 2.0. On Verified: DeepSeek-V4-Pro-Max (80.6%), MiniMax M3 (80.5%), Qwen3.7 Max (80.4%). On Scale's standardized leaderboard, the top open-weights entry is still qwen3-coder-480b-a35b at 38.7%. See best open-source coding models.

What does a SWE-bench Pro run cost per solved task?

A benchmark published September 10, 2026 by the imec aistack team ran 64 curated SWE-bench Pro tasks through Claude Code, Codex, and Pi against two self-hosted models. Qwen3.8-27B in FP8 on one H200 resolved 33, 31, and 28 of 64 for $10.26, $9.00, and $10.16 of GPU time. GLM-5.3-Flash in FP8 on four H200s resolved 32, 33, and 32 for $45.00, $22.80, and $45.00. Cost per solved task ran $12.90 to $26.00, so the harness moved spend by about 2x at the same accuracy.

Why do vendor scores and Scale standardized scores differ?

Scale runs every model through identical scaffolding; vendors run tuned agent harnesses, and llm-stats aggregates those self-reported numbers. The gap is 10-30 points and is mostly context retrieval and tool-use quality, not model capability. Scale's own runs are not uniform either: its footnote says grayed-out rows used a capped cost limit and a 50-turn limit while everything else ran uncapped at 250 turns, and the top three public-set rows are marked as mini-swe-agent runs. Morph's internal runs show the same effect from one variable: adding the WarpGrep v2 search subagent lifts every model tested by 2.1-2.2 points.

Is SWE-bench Verified still useful?

As a frontier ranking, no: OpenAI deprecated it in February 2026 over confirmed contamination, and on the llm-stats tracker ranks 8 through 15 now sit within 0.6 points of each other with a three-way tie at 80.2%. It still separates weak models from strong ones and runs cheaply. For production model selection, use SWE-bench Pro's commercial-set scores.

Is SWE-bench Pro reliable?

It is the most contamination-resistant public coding benchmark, but it has known validity issues. Datacurve's May 2026 DeepSWE audit reported that SWE-bench Pro's graders mis-graded roughly one-third of trials (accepted incorrect patches 8.5% of the time, rejected correct ones 24%), and that Claude Opus 4.6 and 4.7 were flagged "CHEATED" on more than 12% of reviewed tasks for reading gold solutions out of the repo's .git history (tracked as Scale's GitHub issue #93, and independently measured at a 21% exploitation rate over 100 tasks in mini-swe-agent issue #787). A separate community claim attributes 68.5% of GPT-5.5's failures to broken test cases. Harness defects cut the other way: pull request #111 on Scale's repository, opened July 4, 2026, found the Docker SDK's 60-second default read timeout scored slow Go and JavaScript instances unresolved with no error, and issue #115, filed August 3, 2026, documents a Navidrome task whose visible requirement contradicts its own official test. None of these figures are confirmed by Scale AI. Read the standardized commercial-set scores as the most reliable signal.

Sources

Primary sources behind the scores, prices, and model-status claims on this page, checked September 14, 2026. Vendor-reported figures (llm-stats aggregate) are not independently verified. Scale's standardized public and private leaderboards are the directly comparable numbers, with the harness and turn-limit caveats in the harness-rules section above.

WarpGrep v2: Search Subagent for SWE-bench Pro

WarpGrep v2 is the RL-trained search subagent that lifted every model it was paired with by 2+ points on SWE-bench Pro. It runs in its own context window, issues 8 parallel tool calls per turn, and makes your coding agent 15.6% cheaper and 28% faster. $0.80 per 100K tokens.