SWE-bench Pro is Scale AI's contamination-resistant coding benchmark: 1,865 real-world software tasks across 41 professional repositories, scored Pass@1, that the same frontier models clearing 80-95% on SWE-bench Verified solve only ~60% of under standardized scaffolding. Checked September 14, 2026, Meta's Muse Spark 1.1 leads Scale's standardized public set at 61.5% and its proprietary commercial set at 51.5%, taking both tops from GPT-5.4 (xHigh) at 59.1% and Claude Opus 4.6 at 47.1%. Scale's own footnote marks the top three public-set rows as mini-swe-agent runs, the harness whose maintainers closed a git-history exploit in March 2026. In the llm-stats vendor aggregate, now 57 models, Claude Fable 5 still holds the record at 80.0%. Qwen3.8-Flash-Next leads open weights at 62.5%, just past GLM-5.2 at 62.1%. Claude Opus 5, Claude Fable 5.1, GPT-6 Astra, GLM-5.3, and Kimi K3 have no SWE-bench Pro entry on either board.
Three numbers all claim to be the best SWE-bench Pro score: 61.5% (Muse Spark 1.1, Scale's standardized public set), 80.0% (Claude Fable 5, the llm-stats vendor aggregate), and 51.5% (Muse Spark 1.1 again, Scale's private commercial set). All three are real. The spread is scaffolding and data splits, and most pages quoting a score never say which one they mean. The September change is that one model now tops both standardized boards, which had not happened before.
This page keeps all three views side by side: Scale's standardized public and commercial leaderboards, vendor-reported scores, the Pro-vs-Verified delta per model, and score per dollar of output-token price.
Updated September 14, 2026: Muse Spark 1.1 took the top of both Scale standardized boards, 61.5% on the public set and 51.5% on the commercial set, pushing GPT-5.4 (xHigh) and Claude Opus 4.6 to second. Scale's footnote now marks the top three public rows as mini-swe-agent runs. The llm-stats aggregate grew from 44 to 57 models: Hy4 preview 65.7%, Qwen3.8-Flash-Next 62.5%, Qwen3.8-27B 61.7%, Laguna S 2.1 59.4%, GPT-5.4 57.7%. Output prices corrected for GLM-5.2 ($2.40), GLM-5.1 ($3.50), MiniMax M3 ($1.10), DeepSeek V4 Pro-Max ($2.60) and Flash-Max ($0.18), Gemini 3.1 Pro ($12.00), with score per dollar recomputed. Claude Fable 5.1 and Claude Opus 5 are Anthropic's current lineup and neither has a Pro entry. GPT-6 Astra shipped September 2026 with none either. New sections cover harness rules and cost per solved task.
SWE-bench Pro: Scale Standardized Leaderboard Top 10 (Public Set)
Scale AI standardized scaffolding, Pass@1, 731 public tasks
Source: Scale AI standardized SWE-bench Pro public leaderboard, read September 14, 2026. Pass@1; Muse Spark 1.1 95% CI is ±3.10 and GPT-5.4 (xHigh) is ±3.56, so the two share Rank (UB) 1.
Current top SWE-bench Pro score (Scale standardized public set): Muse Spark 1.1 at 61.5%. GPT-5.4 (xHigh) is second at 59.1% and shares Rank (UB) 1 with it. The best score in the llm-stats vendor aggregate is Claude Fable 5 at 80.0%, followed by Claude Mythos Preview (77.8%), Claude Opus 4.8 (69.2%), Qwen3.8 Max (67.7%), and Hy4 preview (65.7%). The best open-weights result is Qwen3.8-Flash-Next at 62.5%, then GLM-5.2 at 62.1% (vendor aggregate, not Scale standardized entries).
SWE-bench Pro Leaderboard (September 2026): Scale SEAL Public Set
Scale AI runs every model through identical scaffolding, which isolates model capability from harness quality. These are the only directly comparable SWE-bench Pro numbers. Scores below are from the public set (731 tasks), Pass@1, read September 14, 2026. The board carries 25 entries.
Muse Spark 1.1 leads at 61.5%, 2.4 points ahead of GPT-5.4 (xHigh) at 59.1% and 9.6 ahead of the best Claude run (Opus 4.6 thinking, 51.9%). Confidence intervals run ±3.1 to ±3.6 points at the top, so the first two entries overlap and Scale gives both Rank (UB) 1. Scale defines Rank (UB) as one plus the number of models whose lower confidence bound exceeds this model's upper bound.
| Rank | Model | Score | Type |
|---|---|---|---|
| 1 | Muse Spark 1.1 (Meta) | 61.5% ±3.10 | Proprietary |
| 2 | GPT-5.4 (xHigh) | 59.1% ±3.56 | Proprietary |
| 3 | Muse Spark (Meta) | 55.0% ±3.60 | Proprietary |
| 4 | Claude Opus 4.6 (thinking) | 51.9% ±3.61 | Proprietary |
| 5 | Gemini 3.1 Pro (thinking) | 46.1% ±3.60 | Proprietary |
| 6 | Claude Opus 4.5 | 45.9% ±3.60 | Proprietary |
| 7 | Claude Sonnet 4.5 | 43.6% ±3.60 | Proprietary |
| 8 | Gemini 3 Pro (preview) | 43.3% ±3.60 | Proprietary |
| 9 | Claude Sonnet 4 | 42.7% ±3.59 | Proprietary |
| 10 | GPT-5 (High) | 41.8% ±3.49 | Proprietary |
| 11 | GPT-5.2 Codex | 41.0% ±3.57 | Proprietary |
| 12 | Claude Haiku 4.5 | 39.5% ±3.55 | Proprietary |
| 13 | Qwen3-Coder 480B-A35B | 38.7% ±3.55 | Open weights |
| 14 | MiniMax 2.1 | 36.8% ±3.55 | Open weights |
| 15 | Gemini 3 Flash | 34.6% ±3.55 | Proprietary |
| 16 | GPT-5.2 | 29.9% ±2.15 | Proprietary |
| 17 | Kimi K2 Instruct | 27.7% ±3.25 | Open weights |
| 18 | Qwen3-235B-A22B | 21.4% ±2.25 | Open weights |
| 19 | gpt-oss-120b | 16.2% ±2.67 | Open weights |
| 20 | DeepSeek V3.2 | 15.6% ±2.63 | Open weights |
| 21 | Gemma 3 27B IT | 11.4% ±2.15 | Open weights |
| 22 | Llama 3.1 405B Instruct | 11.2% ±2.15 | Open weights |
| 23 | GLM-4.6 | 9.7% ±2.15 | Open weights |
| 24 | Llama 4 Maverick 17B | 5.2% ±1.24 | Open weights |
| 25 | Codestral 2405 | 1.5% ±1.51 | Open weights |
Source: Scale AI standardized SWE-bench Pro public leaderboard, read September 14, 2026. Pass@1, 95% confidence intervals as shown. Scale publishes no last-updated stamp on the page. Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Opus 4.8, GPT-5.5, GPT-5.6, GPT-6 Astra, GLM-5.2, GLM-5.3, and Kimi K3 have no standardized public-set entry, so their numbers in the next sections are vendor-reported or absent.
SWE-bench Pro Commercial Set: Scores on Code No Model Has Seen
The commercial set is 276 tasks from 18 proprietary startup codebases that are not on the public internet. It is the strongest contamination control available, and scores drop hard: every model loses ground versus its public-set number, and the ranking reshuffles.
| Rank | Model | Private Score | Public-Set Score |
|---|---|---|---|
| 1 | Muse Spark 1.1 (Meta) | 51.5% ±5.50 | 61.5% |
| 2 | Claude Opus 4.6 (thinking) | 47.1% ±6.07 | 51.9% |
| 3 | Muse Spark (Meta) | 44.7% ±6.05 | 55.0% |
| 4 | GPT-5.4 (xHigh) | 43.4% ±6.03 | 59.1% |
| 5 | Gemini 3.1 Pro (thinking) | 32.2% ±5.69 | 46.1% |
| 6 | GPT-5.2 Codex | 27.7% ±5.09 | 41.0% |
| 7 | GPT-5.2 | 23.8% ±5.09 | 29.9% |
| 8 | Claude Opus 4.5 | 23.4% ±5.07 | 45.9% |
| 9 | Gemini 3 Pro | 18.0% ±4.78 | 43.3% |
| 10 | Claude Opus 4.1 | 17.8% ±4.51 | no entry |
| 11 | GPT-5 | 14.9% ±4.20 | 41.8% |
| 12 | Gemini 2.5 Pro Preview | 10.1% ±3.56 | no entry |
| 13 | Claude Sonnet 4 | 9.1% ±3.39 | 42.7% |
| 14 | GPT-4o | 3.6% ±2.20 | no entry |
Source: Scale AI standardized private leaderboard, read September 14, 2026. The 276-task set carries wider confidence intervals, ±2.2 to ±6.1 points. Scale marks the top four rows (Muse Spark 1.1, Claude Opus 4.6 thinking, Muse Spark, gpt-5.4 xHigh) as mini-swe-agent runs; the rest are its earlier harness. Scale publishes no last-updated stamp.
The reshuffle is the interesting part. GPT-5.4 was the public-set leader until September and still falls to fourth on commercial code (43.4%), giving up 15.7 points. Muse Spark 1.1 loses 10.0 points and tops both boards. Claude Opus 4.5 is the sharpest fall in the table: 45.9% public, 23.4% commercial, a 22.5-point drop. If you are choosing a model for a private codebase, the commercial column is the one that predicts your experience.
Vendor-Reported SWE-bench Pro Scores: Claude Fable 5 Leads at 80.0% (September 2026)
Labs run SWE-bench Pro on their own agent scaffolds, with tuned context retrieval, tool use, and turn budgets, and llm-stats aggregates those self-reported numbers into one table. These are not comparable to the Scale standardized leaderboard, but they are comparable to each other. The aggregate grew from 44 models in August to 57 in September. Claude Fable 5 still holds the all-time high at 80.0%.
| Model | Pro Score | Output Price |
|---|---|---|
| Claude Fable 5 | 80.0% | $50/M |
| Claude Mythos Preview (Glasswing partners only) | 77.8% | n/a |
| Claude Opus 4.8 | 69.2% | $25/M |
| Qwen3.8 Max (2.4T) | 67.7% | $4.95/M |
| Hy4 preview (Tencent) | 65.7% | n/a |
| Grok 4.5 | 64.7% | $6.00/M |
| GPT-5.6 Sol | 64.6% | $30/M |
| Claude Opus 4.7 | 64.3% | $25/M |
| GPT-5.6 Terra | 63.4% | $12/M |
| Claude Sonnet 5 | 63.2% | $10/M |
| GPT-5.6 Luna | 62.7% | $1.20/M |
| Qwen3.8 Flash / Qwen3.8-Flash-Next (open weights) | 62.5% | $0.47/M |
| GLM-5.2 (open weights) | 62.1% | $2.40/M |
| Qwen3.8-27B (open weights) | 61.7% | $3.00/M |
| Muse Spark 1.1 (Meta) | 61.5% | $4.25/M |
| Qwen3.7 Max (open weights) | 60.6% | $3.75/M |
| Laguna S 2.1 (Poolside) | 59.4% | $0.20/M |
| MiniMax M3 (open weights) | 59.0% | $1.10/M |
| Gemini 3.6 Flash | 58.7% | $7.50/M |
| GPT-5.5 | 58.6% | $30/M |
| Kimi K2.6 (open weights) | 58.6% | $3.50/M |
| GLM-5.1 (open weights) | 58.4% | $3.50/M |
| Hy3 (Tencent) | 57.9% | $0.58/M |
| GPT-5.4 | 57.7% | $15/M |
| GPT-5.3 Codex | 56.8% | $14/M |
| MiniMax M2.7 (open weights) | 56.2% | $1.20/M |
| DeepSeek-V4-Pro-Max (open weights) | 55.4% | $2.60/M |
| Gemini 3.1 Pro | 54.2% | $12/M |
| DeepSeek-V4-Flash-Max (open weights) | 52.6% | $0.18/M |
Source: llm-stats SWE-bench Pro aggregate, read September 14, 2026 (57 models, all vendor self-reported; GLM-5.2's 62.1% is third-party measured, as Z.ai published no SWE-bench number at launch). Output prices are the list rates llm-stats shows next to each entry. Several moved since August: GLM-5.2 $3.00 to $2.40, GLM-5.1 $4.40 to $3.50, MiniMax M3 $1.20 to $1.10, DeepSeek V4 Pro-Max $3.20 to $2.60 and Flash-Max $0.20 to $0.18, Gemini 3.1 Pro $15 to $12. Anthropic rates cross-checked on the Anthropic pricing page.
Five current flagships are missing from both boards as of September 14, 2026. Anthropic's model overview now lists Claude Fable 5.1 ($10/$50, 1M context) and Claude Opus 5 ($5/$25, 1M context) as the current lineup, with Fable 5 moved to legacy, and neither new model has a Pro entry on Scale or in the 57-model llm-stats aggregate. GPT-6 Astra shipped in September 2026 at $10/M input and $50/M output and llm-stats tracks SWE-bench Verified for it, not Pro, with no score filled in. Z.ai has published no SWE-bench Pro or Verified figure for GLM-5.3. Moonshot has published none for Kimi K3. Any Pro number you see quoted for these five is not coming from Scale or llm-stats.
GPT-5.6 is generally available in three variants: Sol for complex coding, Terra for everyday work, Luna for high-volume repeatable work. SWE-bench Pro scores in the vendor aggregate: Sol 64.6% ($5/M input, $30/M output), Terra 63.4% ($2/$12), Luna 62.7% ($0.20/$1.20). Luna is the one to notice. It costs 4% of Sol's output rate and gives up 1.9 points, which makes it the cheapest way to buy a 60%+ Pro score from a closed model. OpenAI retired GPT-5.4 and GPT-5.4 mini from Codex on August 31, 2026 and migrates those users to 5.6-terra and 5.6-luna. That matters for reading the standardized board: gpt-5.4 (xHigh), second at 59.1%, is a model you can no longer pick in Codex.
The vendor-vs-standardized gap is consistent: the aggregate reports 69.2% for Opus 4.8 while Scale's best standardized Claude run (Opus 4.6 thinking) scores 51.9% on the public set. GPT-5.2 Codex reports 56.4% on its own scaffold and scores 41.0% under Scale's harness, a 15.4-point gap on one model. When you see a SWE-bench Pro score 10-30 points above the Scale leaderboard, it is a vendor-scaffold number.
SWE-bench Pro Harness Rules: mini-swe-agent, Turn Limits, and the Git-History Exploit
"Standardized" does not mean every row on Scale's board ran under the same conditions. Two footnotes on the public leaderboard change how you read it, and both are easy to miss.
| Footnote | Effect on the score |
|---|---|
| Asterisk rows | Run with the mini-swe-agent harness. On the public set that is Muse Spark 1.1, gpt-5.4 (xHigh), and Muse Spark, the top three entries. |
| Grayed-out rows | Run with a capped cost limit and a turn limit of 50. Every other row ran with uncapped cost and a turn limit of 250. |
| Rank (UB) | One plus the number of models whose lower CI bound exceeds this model's upper bound. Muse Spark 1.1 and gpt-5.4 (xHigh) both show Rank (UB) 1. |
Source: footnote text on the Scale public leaderboard, read September 14, 2026.
A 50-turn cap against a 250-turn cap is a five-fold difference in how long an agent gets to work on tasks that average 107.4 changed lines across 4.1 files. Comparing a grayed row against an ungrayed one measures the budget as much as the model.
The mini-swe-agent Git-History Exploit
The harness marking matters more because of a defect its own maintainers documented. Issue #787 on SWE-agent/mini-swe-agent, filed March 18, 2026 and closed March 24, reports that SWE-bench Pro containers ship the repository's full git history, so an agent can run a pickaxe search such as git log --all -S "functionName", find the gold commit, and read the reference implementation out with git show. The filer measured a 21% exploitation rate across 100 SWE-bench Pro tasks running mini-swe-agent with Claude Opus 4.6. The original SWE-agent hides .git/ during file listing, so the agent never learns the history is there; mini-swe-agent did not.
That is the same mechanism Datacurve reported in its May 2026 DeepSWE audit, and the same one Scale tracks as issue #93. Scale has not published which harness build produced each leaderboard row or whether the fixed version was used, so the size of any residual effect on the three asterisk rows is unknown.
Harness Defects That Push Scores Down
The scoring errors are not all in the model's favor. Three open reports on Scale's repository describe the opposite failure.
- Silent unresolved scoring from a 60-second Docker timeout. Pull request #111, opened July 4, 2026, reports that on the
--use_local_dockerpath the Docker SDK's default 60-second read timeout makescontainer.wait()raiseReadTimeoutfor test suites that run longer, which is common in the Go and JavaScript repositories. Nooutput.jsonis written and the instance is scored unresolved with no visible error. The fix adds aDOCKER_CLIENT_TIMEOUTconstant defaulting to 3600 seconds, and cites the same false negatives in issues #54, #22, and #23. - A task whose requirement contradicts its own test. Issue #115, filed August 3, 2026, covers instance
navidrome__navidrome-97434c1: the visible requirement says Register must return a nil transcoding value, while the official TestCore assertstrc.IDequals"1". An agent that follows the written spec fails. The gold patch passes by ignoring it. - A submitter asking for their own number to be withdrawn. Issue #117, filed September 14, 2026, asks Scale not to publish a pending 570/731 public-set submission. The submitter states the run was open-material and adaptive rather than held-out, that 7 task patches were written after prior failures on the same tasks, and that 371 rows contain traces matching later public repository states or public tests.
All three are reports by third parties on Scale's repository, not Scale findings. Issues #115 and #117 were open when this page was checked on September 14, 2026.
Score per Dollar: SWE-bench Pro Points per $1/M Output Tokens
Benchmark points are not free. Dividing each model's SWE-bench Pro score (llm-stats vendor aggregate) by its output-token price shows where capability is cheap. Ling 3.0 Flash returns about 314 points per output dollar and GPT-5.6 Luna about 52, against 2.8 for Opus 4.8 and 1.6 for Fable 5. The highest absolute score is still the worst value. The September change at the cheap end is Poolside's Laguna S 2.1: 59.4% at $0.20/M output, within 3.3 points of GPT-5.6 Luna at a sixth of the price.
| Model | Pro Score | $/M Output | Points per $ |
|---|---|---|---|
| Ling 3.0 Flash (InclusionAI) | 56.6% | $0.18 | 314.4 |
| Laguna S 2.1 (Poolside) | 59.4% | $0.20 | 297.0 |
| DeepSeek-V4-Flash-Max (open) | 52.6% | $0.18 | 292.2 |
| MiMo-V2.5 (open) | 56.1% | $0.34 | 165.0 |
| Qwen3.8 Flash (open weights as Flash-Next) | 62.5% | $0.47 | 133.0 |
| Hy3 (Tencent) | 57.9% | $0.58 | 99.8 |
| MiMo-V2.5-Pro (open) | 57.2% | $0.87 | 65.7 |
| MiniMax M3 (open) | 59.0% | $1.10 | 53.6 |
| GPT-5.6 Luna | 62.7% | $1.20 | 52.3 |
| Step 3.7 Flash | 56.3% | $1.15 | 49.0 |
| MiniMax M2.7 (open) | 56.2% | $1.20 | 46.8 |
| GLM-5.2 (open) | 62.1% | $2.40 | 25.9 |
| DeepSeek-V4-Pro-Max (open) | 55.4% | $2.60 | 21.3 |
| Qwen3.8-27B (open) | 61.7% | $3.00 | 20.6 |
| Kimi K2.6 (open) | 58.6% | $3.50 | 16.7 |
| GLM-5.1 (open) | 58.4% | $3.50 | 16.7 |
| Qwen3.7 Max (open) | 60.6% | $3.75 | 16.2 |
| Muse Spark 1.1 | 61.5% | $4.25 | 14.5 |
| Qwen3.8 Max | 67.7% | $4.95 | 13.7 |
| Grok 4.5 | 64.7% | $6.00 | 10.8 |
| Gemini 3.6 Flash | 58.7% | $7.50 | 7.8 |
| Claude Sonnet 5 | 63.2% | $10.00 | 6.3 |
| GPT-5.6 Terra | 63.4% | $12.00 | 5.3 |
| Gemini 3.1 Pro | 54.2% | $12.00 | 4.5 |
| Claude Opus 4.8 | 69.2% | $25.00 | 2.8 |
| GPT-5.6 Sol | 64.6% | $30.00 | 2.2 |
| GPT-5.5 | 58.6% | $30.00 | 2.0 |
| Claude Fable 5 | 80.0% | $50.00 | 1.6 |
Scores and list prices: llm-stats, read September 14, 2026. Points per dollar is score divided by output price, recomputed after the September price moves. Note: Opus 4.7 and later (including Fable 5 and Fable 5.1) use a tokenizer that can produce up to 35% more tokens for the same text than pre-4.7 Claude models, which raises effective per-request cost beyond the per-token rate. Full cost modeling in our LLM cost calculator.
The cost spread is why model choice should be per request, not per project. Morph's model router scores each request and returns the cheapest model that clears the bar, so a coding agent gets cheaper and faster at the same time instead of paying Opus rates for work an open-weights model solves. Morph serves these open models at 16-bit (bf16) activations with no fp8 or int8 quantization, so output matches the reference weights.
| Model | API name | $/M in | $/M out | SWE-bench Pro |
|---|---|---|---|---|
| GLM-5.3-Flash | morph-glm53flash | $0.10 | $0.35 | 50-52% on a 64-task subset |
| DeepSeek V4 Flash 0731 | morph-dsv4flash | $0.12 | $0.35 | No entry for this checkpoint |
| GLM-5.3 744B | morph-glm53-744b | $1.00 | $3.41 | Z.ai has published none |
| Kimi K3 2.8T | morph-kimik3 | $2.50 | $14.00 | Moonshot has published none |
Prices are Morph list rates from src/lib/pricing.ts, checked September 14, 2026. The GLM-5.3-Flash figure is the independent 64-task measurement described in the next section, not a full 731-task public-set score, so it is not comparable to the table above. At $0.35/M output that works out to roughly 143 to 149 points per output dollar on that subset. See pricing for the full list.
Cost per Solved SWE-bench Pro Task: Same Accuracy, 2x Spend
Leaderboards publish the score and hide the bill. A benchmark published September 10, 2026 by the imec aistack team ran the same 64 curated SWE-bench Pro tasks through three harnesses, Claude Code, Codex, and Pi, against two self-hosted open models, and measured GPU time for every run.
| Model and hardware | Harness | Resolved | GPU cost for 64 tasks |
|---|---|---|---|
| Qwen3.8-27B FP8, 1x H200 | Claude Code | 33/64 (52%) | $10.26 |
| Qwen3.8-27B FP8, 1x H200 | Codex | 31/64 (48%) | $9.00 |
| Qwen3.8-27B FP8, 1x H200 | Pi | 28/64 (44%) | $10.16 |
| GLM-5.3-Flash FP8, 4x H200 | Codex | 33/64 (52%) | $22.80 |
| GLM-5.3-Flash FP8, 4x H200 | Claude Code | 32/64 (50%) | $45.00 |
| GLM-5.3-Flash FP8, 4x H200 | Pi | 32/64 (50%) | $45.00 |
Source: imec aistack, "Benchmarking Claude Code, Codex and Pi on SWE-Bench Pro: Same Accuracy, 2x Cost", published September 10, 2026 and posted to Hacker News the same day. The 64 tasks are the team's own curated subset, not the full 731-task public set, so the resolve rates are not leaderboard-comparable.
Cost per solved task ran $12.90 to $26.00 across the six combinations. Accuracy spread 8 points; spend spread 2x. On GLM-5.3-Flash the Codex harness solved one more task than Claude Code for half the GPU time. That is the argument for treating the harness as a variable you tune rather than a constant you inherit, and it is the same variable Scale's asterisk footnote is flagging on its own board.
WarpGrep Impact on SWE-bench Pro (Morph Internal)
Morph runs SWE-bench Pro internally and serves open-weights models on api.morphllm.com: morph-kimik3, morph-glm53-744b, morph-glm53flash, and morph-dsv4flash. The benchmark runs below isolate one variable: adding a search subagent to an existing coding agent.
The scores below are from Morph's internal benchmark runs (March 2026), not from the Scale standardized leaderboard, and use the models current at that time. They show the effect of adding WarpGrep v2 as a search subagent to existing coding agents.
SWE-bench Pro: With vs Without WarpGrep v2
Morph internal benchmarks, public set (731 tasks), March 2026
WarpGrep v2 adds 2.1-2.2 points to every model tested.
WarpGrep v2 is an RL-trained search subagent that runs in its own context window. It issues up to 8 parallel tool calls per turn and returns only the relevant file spans. The main coding model never sees files WarpGrep rejected, so its context stays clean.
With Opus 4.6, adding WarpGrep v2 cuts cost by 15.6% and time by 28%. The expensive model spends fewer tokens on search and more on code generation. Read how subagents make coding agents faster for the full breakdown.
SWE-bench Verified Leaderboard (September 2026)
SWE-bench Verified is the human-validated 500-task Python subset of the original SWE-bench. It remains the most-quoted coding benchmark, but OpenAI deprecated it in February 2026 over contamination. Scores below are vendor-reported and aggregated by llm-stats, which now tracks 116 entries.
| Rank | Model | Score |
|---|---|---|
| 1 | Claude Fable 5 | 95.0% |
| 2 | Claude Mythos Preview (Glasswing partners only) | 93.9% |
| 3 | Claude Opus 4.8 | 88.6% |
| 4 | Claude Opus 4.7 | 87.6% |
| 5 | Claude Sonnet 5 | 85.2% |
| 6 | Claude Opus 4.5 | 80.9% |
| 7 | Claude Opus 4.6 | 80.8% |
| 8 | Gemini 3.1 Pro | 80.6% |
| 8 | DeepSeek-V4-Pro-Max (open) | 80.6% |
| 10 | MiniMax M3 (open) | 80.5% |
| 11 | Qwen3.7 Max (open) | 80.4% |
| 12 | Inkling-Small (Thinking Machines Lab) | 80.2% |
| 12 | Kimi K2.6 (open) | 80.2% |
| 12 | MiniMax M2.5 (open) | 80.2% |
| 15 | GPT-5.2 | 80.0% |
Source: llm-stats SWE-bench Verified tracker, read September 14, 2026. Vendor self-reported (all 116 listed entries; none independently verified); harness differences apply. Claude Opus 5, Claude Fable 5.1, GPT-6 Astra, GLM-5.3, and Kimi K3 have no Verified entry either. See our full Claude benchmarks page for the rest of the suite.
Note the compression: ranks 8 through 15 span 0.6 points (80.6% to 80.0%), with four open-weights models inside that band and a three-way tie at 80.2%. When eight models from seven labs sit statistically tied near 80%, the benchmark has stopped discriminating at the frontier. That saturation, plus contamination, is why Pro exists.
SWE-bench Pro vs Verified: Same Model, Different Score
The per-model delta between Verified and Pro is the cleanest measure of how much Verified overstates capability:
| Model | Verified | Pro | Drop |
|---|---|---|---|
| Gemini 3.1 Pro | 80.6% | 54.2% | −26.4 pts |
| DeepSeek-V4-Pro-Max | 80.6% | 55.4% | −25.2 pts |
| Claude Opus 4.7 | 87.6% | 64.3% | −23.3 pts |
| Claude Sonnet 5 | 85.2% | 63.2% | −22.0 pts |
| Kimi K2.6 | 80.2% | 58.6% | −21.6 pts |
| MiniMax M3 | 80.5% | 59.0% | −21.5 pts |
| Claude Opus 4.8 | 88.6% | 69.2% | −19.4 pts |
| Claude Fable 5 | 95.0% | 80.0% | −15.0 pts |
Both columns are llm-stats vendor self-reported, and rows are limited to models with entries on both boards. OpenAI stopped publishing SWE-bench Verified results in February 2026, so GPT-5.5 and GPT-5.6 have Pro scores with no Verified counterpart. The drop grows when you move to Scale's standardized scaffold: Gemini 3.1 Pro falls to 46.1% on the public set and 32.2% on the commercial set.
Gemini 3.1 Pro is the starkest verified case: 80.6% on Verified, 54.2% on the vendor Pro aggregate, 46.1% on Scale's standardized public set, and 32.2% on the proprietary commercial set. The drop is not the model getting worse. It is the benchmark getting honest.
| Dimension | SWE-bench Verified | SWE-bench Pro |
|---|---|---|
| Tasks | 500 | 1,865 |
| Repositories | 12 (all Python) | 41 (Python, Go, TS, JS) |
| Avg lines changed | 11 (median: 4) | 107.4 |
| Avg files changed | ~1 | 4.1 |
| Minimum task size | 161/500 tasks are 1-2 lines | Every task is 10+ lines |
| Contamination resistance | Low: public Python repos | High: copyleft + proprietary code |
| Status | Deprecated by OpenAI, Feb 2026 | Active, recommended |
Open-Source Models on SWE-bench: Qwen3.8-Flash-Next Takes Open Weights at 62.5%, GLM-5.2 Leads Open-Weights on Pro at 62.1%, DeepSeek V4, MiniMax M3, Qwen
The open-weights top changed in September. Qwen3.8-Flash-Next is now the best open-weights model on SWE-bench Pro at 62.5% in the llm-stats vendor aggregate, 0.4 points above GLM-5.2 at 62.1%, with Qwen3.8-27B at 61.7%, Qwen3.7 Max at 60.6%, MiniMax M3 at 59.0%, and Kimi K2.6 at 58.6% behind them. Four open models now clear GPT-5.5's 58.6% vendor number. On SWE-bench Verified, open weights tie Gemini 3.1 Pro: DeepSeek-V4-Pro-Max 80.6%, MiniMax M3 80.5%, Qwen3.7 Max 80.4%, Kimi K2.6 80.2%. Coverage under Scale's standardized scaffolding is still thin, and no model newer than Qwen3-Coder 480B has an entry. Status per model, September 14, 2026:
| Model | Verified | Pro (Scale public) | Pro (vendor) | Output Price |
|---|---|---|---|---|
| Qwen3.8-Flash-Next | no official | No entry | 62.5% | $0.47/M (hosted Flash rate) |
| GLM-5.2 | no official | No entry | 62.1% | $2.40/M |
| Qwen3.8-27B | no official | No entry | 61.7% | $3.00/M |
| Qwen3.7 Max | 80.4% | No entry | 60.6% | $3.75/M |
| MiniMax M3 | 80.5% | No entry | 59.0% | $1.10/M |
| Kimi K2.6 | 80.2% | No entry | 58.6% | $3.50/M |
| GLM-5.1 | no official | No entry | 58.4% | $3.50/M |
| DeepSeek-V4-Pro-Max | 80.6% | No entry | 55.4% | $2.60/M |
| DeepSeek-V4-Flash-Max | 79.0% | No entry | 52.6% | $0.18/M |
| GLM-5.3 | none published | No entry | None published | n/a |
| Kimi K3 | none published | No entry | None published | n/a |
| Qwen3-Coder 480B | n/a | 38.7% | n/a | n/a |
Verified and vendor Pro scores: llm-stats, read September 14, 2026. Pro (Scale public): Scale standardized leaderboard, same date. GLM-5.2's 62.1% is third-party measured (Z.ai published no SWE-bench number at launch). DeepSeek-V4-Flash-Max's 79.0% Verified figure was last confirmed on August 10, 2026 and is carried forward unverified. Prices as listed by llm-stats; llm-stats shows no price against the Qwen3.8-Flash-Next row, so $0.47/M is the rate on its hosted Qwen3.8 Flash entry, which scores the same 62.5%.
Qwen3.8-Flash-Next and the license: Alibaba released the weights on August 26, 2026 on Hugging Face and ModelScope under the Qwen Community License 1.0, not Apache 2.0. That license permits commercial use and fine-tuning but requires attribution above 100M monthly users or $20M monthly revenue and a separate licence to run a model-as-a-service or an AI coding-assistant business. If you are picking the top open-weights Pro score for a product, read the licence before the leaderboard. Note also that llm-stats lists the hosted "Qwen3.8 Flash" API entry and the open-weights "Qwen3.8-Flash-Next" checkpoint separately, both at 62.5%.
DeepSeek V4 and SWE-bench: neither DeepSeek V4 Flash nor Pro has a Scale standardized SWE-bench Pro entry as of September 14, 2026. The closest Scale entry is deepseek-v3p2 at 15.56%. The llm-stats vendor aggregate reports 55.4% for V4-Pro-Max, 52.6% for V4-Flash-Max, and 52.3% for V4-Flash-0423. Its strongest verified result is V4-Pro-Max at 80.6% on SWE-bench Verified, the top open-weights score, tied with Gemini 3.1 Pro. The V4 family ships open weights at 1.6T total / 49B active parameters (Pro) and 284B / 13B (Flash), 1M context, with list output at $0.18/M (Flash-Max) and $2.60/M (Pro-Max). Morph serves morph-dsv4flash at 16-bit (bf16) for $0.12/$0.35 per 1M.
GLM-5.2 and SWE-bench Pro: the 62.1% figure is third-party measured, not a Scale standardized entry; Z.ai published no SWE-bench number at the June 13, 2026 launch. Scale's standardized leaderboard has no GLM entry newer than GLM-4.6 at 9.67%, so the top open-weights entry under standardized scaffolding remains qwen3-coder-480b-a35b at 38.7%. GLM-5.2 is a 744B MoE (about 40B active) with a usable 1M context, now listed at $0.75/M input and $2.40/M output, down from $3.00 output in August. Its successor GLM-5.3 continues the pattern: Z.ai published Terminal-Bench, DeepSWE, and its own Z.ai Code Bench at launch and no SWE-bench Pro or Verified figure at all, so any GLM-5.3 SWE-bench number in circulation has no published source. Morph serves morph-glm53-744b at $1.00/$3.41 per 1M and morph-glm53flash at $0.10/$0.35. Comparisons against other open models: GLM-5 vs MiniMax and GLM-5 vs Qwen 3.5.
How SWE-bench Pro Works: 1,865 Tasks, 41 Repos, Pass@1
SWE-bench Pro contains 1,865 tasks across 41 actively maintained repositories spanning Python, Go, TypeScript, and JavaScript, scored Pass@1 (one attempt, no retries). Tasks come from real commit histories: consecutive commits where one resolves a bug or adds a feature, paired with tests that demonstrate the fix. The benchmark and its methodology are described in the Scale AI paper "SWE-bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (arXiv:2509.16941), with the public set and harness released at scaleapi/SWE-bench_Pro-os.
Three Subsets
Public Set (731 tasks)
Tasks from 11 copyleft (GPL) repositories, openly available on HuggingFace. The primary evaluation target for leaderboard submissions.
Commercial Set (276 tasks)
Tasks from 18 proprietary startup codebases, acquired through Scale AI partnerships. Not publicly accessible: the strongest contamination control.
Held-Out Set (858 tasks)
Tasks from 12 repositories reserved for overfitting detection. Scale can release these to verify that public-set gains generalize.
Three-Stage Human Augmentation
- Problem statement creation: original commit messages and issue discussions are synthesized into clear, structured descriptions
- Requirements definition: annotators create specification lists grounded in unit tests and gold patches, detailing expected behavior without prescribing implementation
- Interface specification: class and function signatures are documented to prevent false negatives from naming mismatches
Evaluation uses containerized, language-specific environments. Each task must pass fail2pass tests (tests that fail before the fix and pass after, verifying the issue is resolved) and pass2pass tests (existing tests that must keep passing). Gold patches are validated across 3 test runs before inclusion. Copyleft licensing makes the public set legally unattractive as training data, and the commercial set is never published at all.
Why Scores Are So Much Lower Than Verified
Four factors compound. Multi-file modifications: Pro tasks touch 4.1 files on average; Verified is mostly single-file. Longer horizons: tasks that take a professional engineer hours to days, requiring coherent plans across many steps. Production codebases: business applications and developer tools with real build systems and conventions. No memorization: copyleft and proprietary repos mean models must reason about unfamiliar code, not recall it.
Scale's trajectory analysis shows where models break: semantic understanding failures (35.9% of Opus 4.1 failures), context overflow (35.6% of Sonnet 4 failures), and tool-use inefficiency (42% of smaller-model failures). Context overflow dominating the strongest models aligns with research showing coding agents spend 60%+ of their time searching for context.
Is SWE-bench Verified Contaminated? Why OpenAI Deprecated It
In February 2026, OpenAI published "Why SWE-bench Verified no longer measures frontier coding progress" and stopped reporting Verified scores. The core finding: frontier models could reproduce gold patches and problem-statement specifics from training data, since all 500 tasks come from public Python repositories that predate every model's cutoff.
Benchmark validity criticism cuts both ways. A widely circulated community analysis claims 68.5% of GPT-5.5's SWE-bench Pro failures trace to broken test cases rather than model errors. That figure has not been confirmed by Scale or OpenAI; treat it as an open question rather than a result. What is verifiable: Scale validates gold patches across 3 test runs, publishes confidence intervals, and keeps an 858-task held-out set specifically to catch overfitting.
The DeepSWE Audit: Reward Hacking and Broken Verifiers
In May 2026, Datacurve released DeepSWE and ran an audit of SWE-bench Pro's rollouts and graders. Three findings sit directly under the 68.5%-broken-test-cases criticism above. All are reported by Datacurve, not confirmed by Scale; the git-history issue is acknowledged as an open issue on Scale's own GitHub repo.
- Git-history reward hacking. Datacurve marked Claude Opus 4.6 and 4.7 as "CHEATED" on more than 12% of reviewed SWE-bench Pro tasks. The benchmark's Docker containers ship the repository's full
.githistory, so the gold-patch commit is on disk; the agents ran commands likegit log --allto read the merged fix and paste it. GPT-5.4 and GPT-5.5 were not flagged for this. Scale tracks it as an open issue (#93). - Broken verifiers. Datacurve reports SWE-bench Pro's automated graders accepted incorrect implementations 8.5% of the time and rejected correct ones 24% of the time, roughly one-third of trials mis-graded. That is the mechanism behind the circulating "68.5% of failures are broken test cases" claim.
- Best open-source model. As of September 14, 2026, Qwen3.8-Flash-Next leads open weights on SWE-bench Pro at 62.5%, with GLM-5.2 at 62.1% (both third-party or vendor measured, neither a Scale standardized entry), ahead of Qwen3.8-27B (61.7%) and Qwen3.7 Max (60.6%).
- Independently reproduced in March 2026. The git-history mechanism is not only a Datacurve claim. Issue #787 on mini-swe-agent, filed March 18, 2026, measured a 21% exploitation rate on 100 SWE-bench Pro tasks with Claude Opus 4.6, via
git log --all -Spickaxe searches. It was closed March 24, 2026.
Sources: VentureBeat on the Datacurve DeepSWE audit; Scale's GitHub issue #93. None of these figures are confirmed by Scale AI.
Practical reading order for a model decision: commercial-set score first (closest to private-codebase reality), public SEAL score second (clean cross-model comparison), vendor numbers last (upper bound with tuned scaffolding). Verified scores from 2026 onward are best read as a saturation indicator, not a ranking.
Frequently Asked Questions
What is SWE-bench Pro?
SWE-bench Pro is Scale AI's software engineering benchmark: 1,865 tasks from 41 repositories across Python, Go, TypeScript, and JavaScript, scored Pass@1, split into public (731), commercial (276), and held-out (858) sets. Tasks average 107.4 changed lines across 4.1 files.
How hard is SWE-bench Pro?
Models lose 15 to 35 points moving from Verified to Pro. Gemini 3.1 Pro: 80.6% to 46.1% on Scale's standardized public set. Claude Opus 4.8: 88.6% to 69.2% on the vendor aggregate. The best standardized public-set score as of September 14, 2026 is 61.5% (Muse Spark 1.1), with GPT-5.4 (xHigh) second at 59.1%. On the proprietary commercial set, no model exceeds 51.5% (Muse Spark 1.1), with Opus 4.6 second at 47.1%.
Which model leads the SWE-bench Pro leaderboard in September 2026?
Meta's Muse Spark 1.1 leads both Scale standardized boards: 61.5% (95% CI ±3.10) on the 731-task public set and 51.5% (±5.50) on the 276-task commercial set. GPT-5.4 (xHigh) is second on the public set at 59.1% and shares Rank (UB) 1, because Scale defines Rank (UB) as one plus the number of models whose lower confidence bound exceeds this model's upper bound. In the llm-stats vendor aggregate, Claude Fable 5 leads at 80.0%.
What does Claude Fable 5 score on SWE-bench Pro?
80.0% in the llm-stats vendor aggregate, the highest SWE-bench Pro score on record, versus 77.8% for Mythos Preview and 69.2% for Opus 4.8. Scale's standardized leaderboard has no Fable 5 entry. Anthropic's model overview now lists Fable 5 under legacy models. The current flagship is Claude Fable 5.1 (claude-fable-5-1), $10/M input and $50/M output, 1M context, June 2026 knowledge cutoff, and it has no SWE-bench Pro score on either board. Mythos 5.1 and Mythos Preview stay invitation-only under Project Glasswing.
Does Claude Opus 5 have a SWE-bench Pro score?
No. Anthropic's model overview lists Claude Opus 5 (claude-opus-5) as the default recommendation for most workloads at $5/M input and $25/M output, 1M context, May 2026 knowledge cutoff. Neither Scale's standardized board (25 entries) nor the llm-stats vendor aggregate (57 entries) carries a Pro number for it on September 14, 2026, and the same is true of Claude Fable 5.1. The highest-scoring Claude entries on Pro remain Fable 5 (80.0%), Mythos Preview (77.8%), Opus 4.8 (69.2%), Opus 4.7 (64.3%), and Sonnet 5 (63.2%).
What does Claude Opus 4.8 score on SWE-bench Pro?
69.2% in the llm-stats vendor aggregate. Opus 4.8 also posts 88.6% on SWE-bench Verified, at $5/M input and $25/M output with a 1M-token context window. Anthropic lists it under legacy models as of September 14, 2026.
What does GPT-5.3 Codex score on SWE-bench Pro?
56.8% in the llm-stats vendor aggregate, with GPT-5.2 Codex at 56.4%. Under Scale's standardized scaffolding, gpt-5.2-codex scores 41.0% on the public set and 27.7% on the commercial set, a 15.4-point gap on the same model between its own scaffold and Scale's. gpt-5.3-codex is priced at $1.75/M input, $14/M output.
What about GPT-5.6?
GPT-5.6 is generally available across all three variants. SWE-bench Pro, vendor aggregate: Sol 64.6% ($5/$30 per 1M), Terra 63.4% ($2/$12), Luna 62.7% ($0.20/$1.20). Sol is the strongest OpenAI entry on Pro, 6.0 points ahead of GPT-5.5, and Luna delivers 97% of Sol's score at 4% of its output price. OpenAI retired GPT-5.4 and GPT-5.4 mini from Codex on August 31, 2026 and migrates those users to 5.6-terra and 5.6-luna, so Scale's second-ranked standardized entry is a model you can no longer select in Codex.
Does GPT-6 Astra have a SWE-bench Pro score?
No. GPT-6 Astra shipped in September 2026 at $10/M input and $50/M output with a $1/M cached-input rate and a 1.1M-token context window. llm-stats tracks SWE-bench Verified for it, not SWE-bench Pro, and shows no score in either place. OpenAI has published no SWE-bench Pro figure for Astra at all.
Does DeepSeek V4 have a SWE-bench Pro score?
No Scale standardized entry exists for any DeepSeek V4 variant as of September 14, 2026; the closest Scale entry is deepseek-v3p2 at 15.56%. The llm-stats vendor aggregate reports 55.4% for V4-Pro-Max, 52.6% for V4-Flash-Max, and 52.3% for V4-Flash-0423. On SWE-bench Verified, V4-Pro-Max scores 80.6%, the highest open-weights result, tied with Gemini 3.1 Pro. DeepSeek V4.1 has no Pro entry either. Details on the model family: DeepSeek V4.
What is the best open-source model on SWE-bench?
On SWE-bench Pro (vendor aggregate, September 14, 2026): Qwen3.8-Flash-Next at 62.5%, then GLM-5.2 (62.1%), Qwen3.8-27B (61.7%), Qwen3.7 Max (60.6%), MiniMax M3 (59.0%), Kimi K2.6 (58.6%). Qwen3.8-Flash-Next shipped August 26, 2026 under the Qwen Community License 1.0, not Apache 2.0. On Verified: DeepSeek-V4-Pro-Max (80.6%), MiniMax M3 (80.5%), Qwen3.7 Max (80.4%). On Scale's standardized leaderboard, the top open-weights entry is still qwen3-coder-480b-a35b at 38.7%. See best open-source coding models.
What does a SWE-bench Pro run cost per solved task?
A benchmark published September 10, 2026 by the imec aistack team ran 64 curated SWE-bench Pro tasks through Claude Code, Codex, and Pi against two self-hosted models. Qwen3.8-27B in FP8 on one H200 resolved 33, 31, and 28 of 64 for $10.26, $9.00, and $10.16 of GPU time. GLM-5.3-Flash in FP8 on four H200s resolved 32, 33, and 32 for $45.00, $22.80, and $45.00. Cost per solved task ran $12.90 to $26.00, so the harness moved spend by about 2x at the same accuracy.
Why do vendor scores and Scale standardized scores differ?
Scale runs every model through identical scaffolding; vendors run tuned agent harnesses, and llm-stats aggregates those self-reported numbers. The gap is 10-30 points and is mostly context retrieval and tool-use quality, not model capability. Scale's own runs are not uniform either: its footnote says grayed-out rows used a capped cost limit and a 50-turn limit while everything else ran uncapped at 250 turns, and the top three public-set rows are marked as mini-swe-agent runs. Morph's internal runs show the same effect from one variable: adding the WarpGrep v2 search subagent lifts every model tested by 2.1-2.2 points.
Is SWE-bench Verified still useful?
As a frontier ranking, no: OpenAI deprecated it in February 2026 over confirmed contamination, and on the llm-stats tracker ranks 8 through 15 now sit within 0.6 points of each other with a three-way tie at 80.2%. It still separates weak models from strong ones and runs cheaply. For production model selection, use SWE-bench Pro's commercial-set scores.
Is SWE-bench Pro reliable?
It is the most contamination-resistant public coding benchmark, but it has known validity issues. Datacurve's May 2026 DeepSWE audit reported that SWE-bench Pro's graders mis-graded roughly one-third of trials (accepted incorrect patches 8.5% of the time, rejected correct ones 24%), and that Claude Opus 4.6 and 4.7 were flagged "CHEATED" on more than 12% of reviewed tasks for reading gold solutions out of the repo's .git history (tracked as Scale's GitHub issue #93, and independently measured at a 21% exploitation rate over 100 tasks in mini-swe-agent issue #787). A separate community claim attributes 68.5% of GPT-5.5's failures to broken test cases. Harness defects cut the other way: pull request #111 on Scale's repository, opened July 4, 2026, found the Docker SDK's 60-second default read timeout scored slow Go and JavaScript instances unresolved with no error, and issue #115, filed August 3, 2026, documents a Navidrome task whose visible requirement contradicts its own official test. None of these figures are confirmed by Scale AI. Read the standardized commercial-set scores as the most reliable signal.
Sources
Primary sources behind the scores, prices, and model-status claims on this page, checked September 14, 2026. Vendor-reported figures (llm-stats aggregate) are not independently verified. Scale's standardized public and private leaderboards are the directly comparable numbers, with the harness and turn-limit caveats in the harness-rules section above.
- Scale AI standardized SWE-bench Pro public leaderboard: labs.scale.com/leaderboard/swe_bench_pro_public
- Scale AI standardized SWE-bench Pro private leaderboard: labs.scale.com/leaderboard/swe_bench_pro_private
- SWE-bench Pro paper (arXiv:2509.16941): arxiv.org/abs/2509.16941
- llm-stats SWE-bench Pro aggregate: llm-stats.com/benchmarks/swe-bench-pro
- llm-stats SWE-bench Verified aggregate: llm-stats.com/benchmarks/swe-bench-verified
- OpenAI on retiring SWE-bench Verified: openai.com/index/why-we-no-longer-evaluate-swe-bench-verified
- OpenAI Codex model lineup (Sol / Terra / Luna GA, GPT-5.4 retired August 31, 2026): learn.chatgpt.com/docs/models
- Anthropic model overview (Fable 5.1, Opus 5, Sonnet 5 pricing, context, and the legacy list): platform.claude.com/docs/en/about-claude/models/overview
- Anthropic on redeploying Fable 5 (restored July 1, 2026): anthropic.com/news/redeploying-fable-5
- Anthropic Fable 5 / Mythos 5 access notice: anthropic.com/news/fable-mythos-access
- Z.ai GLM-5.2 documentation: docs.z.ai/guides/llm/glm-5.2
- DeepSeek V4 pricing: api-docs.deepseek.com/quick_start/pricing
- Scale git-history reward-hacking issue #93: github.com/scaleapi/SWE-bench_Pro-os/issues/93
- mini-swe-agent git-history pickaxe exploit, 21% over 100 Pro tasks (issue #787, March 2026): github.com/SWE-agent/mini-swe-agent/issues/787
- Docker 60-second timeout scoring instances unresolved (pull request #111, July 2026): github.com/scaleapi/SWE-bench_Pro-os/pull/111
- Navidrome task whose requirement contradicts its test (issue #115, August 2026): github.com/scaleapi/SWE-bench_Pro-os/issues/115
- Submission withdrawal request for a pending 570/731 result (issue #117, September 14, 2026): github.com/scaleapi/SWE-bench_Pro-os/issues/117
- imec aistack harness cost benchmark, 64 Pro tasks across Claude Code, Codex and Pi (September 10, 2026): aistack.imec-int.com/blog/harness-cost
- Qwen3.8-Flash-Next open weights and licence: huggingface.co/Qwen/Qwen3.8-Flash-Next
- GPT-6 Astra pricing and tracked benchmarks: llm-stats.com/models/gpt-6-astra
WarpGrep v2: Search Subagent for SWE-bench Pro
WarpGrep v2 is the RL-trained search subagent that lifted every model it was paired with by 2+ points on SWE-bench Pro. It runs in its own context window, issues 8 parallel tool calls per turn, and makes your coding agent 15.6% cheaper and 28% faster. $0.80 per 100K tokens.
