Best AI Model for Coding: Quick Answer (July 2026)
The best AI model for coding in July 2026 depends on the job. Top raw capability: GPT-5.6 Sol (96.2% SWE-bench Verified, Vals AI independent), then Claude Fable 5 (95.0%) and Kimi K3 (93.4%). Best for frontend and UI: Kimi K3, the first open model to top the Arena.ai Frontend Code Arena, ahead of Fable 5. Everyday default: Claude Opus 4.8 (88.6%, $5/$25). Best open-weight value: GLM-5.2 ($1.10/$4.10 on Morph), top open model on the Artificial Analysis Intelligence Index.
"Best AI model for coding" and "best LLM for coding" are the same question, and this page answers it by cost per completed task. It divides the two columns every other ranking keeps apart, and adds the ones almost no one measures: an independent SWE-bench Verified harness (Vals AI), the Arena.ai Frontend Code Arena, Scale's standardized SWE-bench Pro leaderboard, official per-token prices, and the output-dollar cost per solved benchmark point. For the same models ranked picks-by-job rather than by cost, see best LLM for coding. Updated July 24, 2026: GPT-5.6 Sol reached general availability July 9 and leads Vals AI at 96.2%, Kimi K3 took #1 on the Frontend Code Arena, GLM-5.2 became the top open model on the Artificial Analysis Intelligence Index, and DeepSeek V4 went GA on July 19.
Highest score
GPT-5.6 Sol
- 96.2% SWE-bench Verified (Vals AI)
- $5 / $30 per M tokens, ~1M context
- GA July 9; Fable 5 (95.0%) and Kimi K3 (93.4%) close behind
Best for frontend / UI
Kimi K3
- #1 Arena.ai Frontend Code Arena (~1,679)
- Ahead of Claude Fable 5, first open model to lead
- 93.4% SWE-bench Verified (Vals AI); weights due ~Jul 27
Best open-weight value
GLM-5.2
- #1 open model, Artificial Analysis Intelligence Index
- $1.10 / $4.10 on Morph (Z.ai lists $1.40 / $4.40)
- 744B MoE / 40B active, MIT, 1M context
GPT-5.6 (Sol, Terra, Luna) cleared its safety review and reached general availability on July 9, 2026; Sol tops Vals AI's independent SWE-bench Verified harness at 96.2% ($5/$30). Kimi K3 launched July 16 and is #1 on the Arena.ai Frontend Code Arena, ahead of Claude Fable 5, with open weights due on Hugging Face around July 27. Claude Fable 5 remains live on the API (restored July 1) at 95.0% Verified. DeepSeek V4 reached general availability July 19 (V4 Pro and V4 Flash, both MIT). GLM-5.2 (744B MoE / 40B active) is the top open model on the Artificial Analysis Intelligence Index, served on Morph at $1.10/$4.10.
13 Models Ranked: SWE-bench Pro x Price x Cost per Solved Point
On a standardized harness, Muse Spark 1.1 solves the most SWE-bench Pro tasks (61.50%), gpt-5.4 follows at 59.10%, and Claude Haiku 4.5 solves them for the least money (about $0.13 of output per point). The table below uses Scale's SEAL public-set leaderboard, which runs every model through the same scaffolding on SWE-bench Pro (1,865 tasks, 41 professional repositories, scored Pass@1). The newest frontier models (GPT-5.6, Fable 5, Opus 4.8, Kimi K3) are not on Scale's standardized board yet, so they appear on the independent SWE-bench Verified section instead. The last column divides official output price by score: dollars of output tokens per benchmark point solved. Lower is more cost-effective.
| Model | SWE-bench Pro | $/M input / output | Output $ per Pro point |
|---|---|---|---|
| Muse Spark 1.1 (Meta) | 61.50% | not published | n/a |
| gpt-5.4 (xHigh) | 59.10% | $2.50 / $15 | $0.25 |
| Muse Spark (Meta) | 55.00% | not published | n/a |
| Claude Opus 4.6 (thinking) | 51.90% | $5 / $25 | $0.48 |
| Gemini 3.1 Pro (thinking) | 46.10% | $2 / $12 (≤200K tokens) | $0.26 |
| Gemini 3 Pro (preview) | 43.30% | preview | n/a |
| gpt-5.2-codex | 41.04% | superseded by gpt-5.5 | n/a |
| Claude Haiku 4.5 | 39.45% | $1 / $5 | $0.13 |
| Qwen3 Coder 480B (open weights) | 38.70% | self-host | n/a |
| Gemini 3 Flash | 34.63% | see Gemini pricing | n/a |
| Kimi K2 Instruct (open weights) | 27.67% | self-host | n/a |
Three things fall out of the combined view. Muse Spark 1.1 and gpt-5.4 top the standardized set, and gpt-5.4 is competitive on cost per point at $0.25. Haiku 4.5 solves about two thirds as many tasks as gpt-5.4 at a third of its output price, making it the cost-per-point leader at roughly $0.13. And Opus 4.6, the top Claude entry Scale has tested, pays a 2x cost-per-point premium over gpt-5.4 for 7.2 fewer points on this harness, which is exactly why Anthropic publishes its own numbers (covered below).
Scale also runs a private (commercial) set drawn from proprietary startup codebases, 276 instances the models have never seen. Models that top the public set do not automatically top unseen code: Muse Spark, for example, scores 55.00% on the public set but 44.70% on the private set. If your repo looks nothing like open-source Python, weight the private-set ordering, not the public leaderboard.
Frontend Code Arena: Kimi K3 Is #1, Ahead of Fable 5
SWE-bench measures repo-scale bug fixing on Python. It says almost nothing about the work most developers do daily: building UI. The Arena.ai Frontend Code Arena fills that gap with blind pairwise human votes on real frontend prompts, and in July 2026 the model on top is Kimi K3, the first open model to lead it, ahead of Claude Fable 5.
Kimi K3 (Moonshot AI) launched July 16, 2026 and took #1 on the Frontend Code Arena at roughly 1,679 points, with Claude Fable 5 second and GPT-5.6 Sol third. It is the first open model to top a frontend coding leaderboard, and on Vals AI's independent SWE-bench Verified harness it scores 93.4%, third behind GPT-5.6 Sol (96.2%) and Fable 5 (95.0%). The catch is serving: on Moonshot's own API, Artificial Analysis measures Kimi K3 at about 34 tokens per second with a 7-second time to first token, and the open weights are not out yet (a Hugging Face release is expected around July 27), so no third party can self-host it faster. That is the gap the Morph Kimi K3 API fills, serving the same model at about 100 tokens per second on an OpenAI-compatible endpoint.
Kimi K3 is #1 for frontend but third on general SWE-bench Verified; GPT-5.6 Sol is first on Verified but third on the Frontend Code Arena. There is no single "best coding model" across both. Pick the leaderboard that matches your work: the Frontend Code Arena for UI-heavy building, the SWE-bench Verified board below for repo-scale bug fixing.
Cost per Completed Task, Not per Token
The per-token price on a model card is close to useless for budgeting a coding agent. What you pay is the price of every token the model burns to finish a task: reasoning tokens, retries, and tool-call round trips. A model that is 12x cheaper per token can end up more expensive per task if it thinks 20x longer, and a model that is pricier per token can be cheaper per task if it solves in one pass. Rank by cost per completed task on your own traffic, not by the headline rate.
Artificial Analysis now scores coding agents on exactly this: average pay-per-token API cost per task, counting input, cache, reasoning, and output tokens separately (Coding Agent Index methodology). The gap this exposes is large. In one worked comparison, Cursor Composer costs $2.50/M output against GPT-5.5's $30/M, a 12x per-token difference, but averages about $0.07 per coding task versus $4.82, a 69x per-task difference, because the workhorse model burns far fewer tokens to close each task (UsageBox, 2026).
Two rules fall out of this for picking a coding LLM. First, a reasoning-heavy frontier model with a low per-token headline can still be the most expensive line on your bill, so measure tokens-per-task before you commit. Second, the cheapest per-token open models (DeepSeek V4 Flash at $0.14/$0.28, MiniMax M3 at $0.60/$2.40) only win on cost per task if they also solve in comparable token counts, which is why a difficulty router that sends the easy 80% to a cheap model and reserves a frontier model for the hard 20% beats any single-model default. The routing section below has the split.
SWE-bench Verified Leaderboard (July 2026)
On Vals AI's independent SWE-bench Verified harness, GPT-5.6 Sol leads at 96.2%, then Claude Fable 5 (95.0%) and Kimi K3 (93.4%). Vals AI runs a single standardized harness across models, which makes it the cleaner Verified reference: the broader llm-stats board is entirely vendor self-reported (0 of 104 entries independently verified). SWE-bench Verified is older, Python-only, and partially contaminated, but it is still the number every launch post quotes. The independent top tier, then the wider vendor board:
SWE-bench Verified: Independent Harness (July 2026)
Source: Vals AI (vals.ai/benchmarks/swebench), single standardized harness. Higher = more GitHub issues resolved.
Vals AI runs one harness across models; GPT-5.6 Sol edges Fable 5, and Kimi K3 sits third while leading the Frontend Code Arena.
The wider vendor-reported board (llm-stats, July 2026) fills in the tier below the frontier: Claude Opus 4.7 (87.6%), Claude Sonnet 5 (85.2%), then the open-weight 80-percent cluster: DeepSeek-V4-Pro-Max (80.6%), Gemini 3.1 Pro (80.6%), MiniMax M3 (80.5%), Qwen3.7 Max (80.4%), and Kimi K2.6 (80.2%). Two structural facts hold. First, the very top is a three-way race now: GPT-5.6 Sol, Fable 5, and Kimi K3 sit within 2.8 points on the independent harness, and Opus 4.8 at $5/$25 covers most teams below them. Second, the 80-percent cluster is mostly open weights; DeepSeek-V4-Pro-Max ties Gemini 3.1 Pro exactly, and you can download its MIT-licensed weights.
Which Claude Model Is Best for Coding?
Claude Fable 5 (claude-fable-5) is the highest-scoring Claude for coding at 95.0% SWE-bench Verified and is live on the API (restored July 1 after a brief export-control suspension), but at $10/$50 it is twice the price of the everyday default, Claude Opus 4.8 (claude-opus-4-8): 88.6% SWE-bench Verified, 69.2% SWE-bench Pro on Anthropic's harness, $5/$25 per million tokens, 1M context with no long-context surcharge. Anthropic ships five current coding-relevant models. Exact API IDs, prices, and the decision logic:
| Model (API ID) | Coding benchmarks | $/M in / out | Context / max output |
|---|---|---|---|
| Claude Fable 5 (claude-fable-5) | 95.0% SWE-bench Verified, 80.0% SWE-bench Pro (vendor); live on API | $10 / $50 | 1M / 128K |
| Claude Opus 4.8 (claude-opus-4-8) | 88.6% Verified, 69.2% Pro (vendor) | $5 / $25 | 1M / 128K |
| Claude Opus 4.7 (claude-opus-4-7) | 87.6% Verified | $5 / $25 | 1M / 128K |
| Claude Sonnet 4.6 (claude-sonnet-4-6) | 79.6% SWE-bench Verified | $3 / $15 | 1M / 64K |
| Claude Haiku 4.5 (claude-haiku-4-5) | 39.45% SWE-bench Pro (Scale SEAL) | $1 / $5 | 200K / 64K |
Default: Opus 4.8
claude-opus-4-8 at $5/$25 is the working default for coding agents: 88.6% SWE-bench Verified, 69.2% SWE-bench Pro on Anthropic's harness (the highest of any buyable model), and 1M context with no long-context surcharge. A fast-mode research preview is priced at $10/$50, about 3x cheaper than fast mode on Opus 4.6/4.7.
Ceiling: Fable 5
claude-fable-5 ($10/$50) adds 6.4 points of SWE-bench Verified and 10.8 points of vendor SWE-bench Pro over Opus 4.8. Suspended June 12 under a US export-control directive, then restored July 1, 2026 after the order was lifted, it is live on the API and the top-scoring Claude; Opus 4.8 at half the price covers most teams.
Volume: Sonnet 4.6
claude-sonnet-4-6 at $3/$15 carries a 1M context window and scores 79.6% SWE-bench Verified. Use it for high-throughput agent loops where Opus pricing compounds: CI review bots, test generation, batch transforms. Batch API halves it to $1.50/$7.50.
Quick edits and subagents: Haiku 4.5
claude-haiku-4-5 at $1/$5 is the cost-per-point leader on Scale's leaderboard (~$0.13 of output per Pro point). Route single-file edits, lint fixes, and explore-style subagents here; cache hits cost $0.10/M.
Treat Claude Sonnet 4, Opus 4, and Opus 4.1 as legacy and migrate to claude-sonnet-4-6 / claude-opus-4-8. Note the tokenizer change too: Opus 4.7 and later (including Fable 5) can produce up to 35% more tokens for the same text than pre-4.7 models, so compare per-request costs, not just per-token rates. Full price tables on the Anthropic API pricing page.
One model to set aside: Claude Mythos 5 (93.9% SWE-bench Verified as Mythos Preview) is a limited-availability model restricted to approved Project Glasswing partners. Unlike Fable 5, which returned to the API on July 1, Mythos 5 remains limited to approved partners. There is no self-serve access, so it is not a practical coding pick.
Claude Opus 4.8 vs GPT-5.6 Sol: The Everyday Frontier Pair
GPT-5.6 Sol reached general availability on July 9, 2026 and is now OpenAI's top model on the API and through Codex. On Vals AI's independent SWE-bench Verified harness it leads at 96.2%, well ahead of Claude Opus 4.8 at 88.6%. Opus 4.8 answers on price and repo-scale engineering: it costs less output ($25 vs $30 per M) and leads SWE-bench Pro on vendor harnesses (69.2% vs Sol's 64.6%). The splits:
| Dimension | Claude Opus 4.8 | GPT-5.6 Sol |
|---|---|---|
| SWE-bench Verified (Vals AI, independent) | 88.6% | 96.2% |
| SWE-bench Pro (vendor) | 69.2% | 64.6% |
| Pricing ($/M in / out) | $5 / $25 | $5 / $30 |
| Cached input ($/M) | $0.50 (cache hit) | $0.50 |
| Context window | 1M | ~1M (128K max output) |
Head-to-Head: The Race Card
GPT-5.3 Codex vs Claude Opus 4.6 across 7 dimensions
Scores based on benchmarks, developer surveys, and hands-on testing as of February 2026. Neither model "wins" overall — it depends on your workflow.
GPT-5.6 Sol is the raw-capability leader on the independent Verified harness and is what OpenAI now points Codex at: the dedicated -codex variants are retired, so the general gpt-5.6 is the current Codex model. Opus 4.8 wins repo-scale software engineering on vendor SWE-bench Pro (69.2% vs 64.6%) and is cheaper on output. If your work lives in a CLI agent, Codex pricing changes the math; if it lives in long-horizon repo edits at half the output price, Opus 4.8 does.
OpenAI previewed GPT-5.6 on June 26, 2026 in three variants: Sol (flagship), Terra (balanced), Luna (fast). It launched as a government-gated limited preview, cleared its cybersecurity safety review early, and reached general availability on July 9, 2026 across ChatGPT, the API, Codex, and GitHub Copilot. Sol is $5/$30 (Terra $2.50/$15, Luna $1/$6) and leads Vals AI's SWE-bench Verified at 96.2%. GPT-5.5 ($5/$30) remains available for teams already on it.
Open-Source Models: 80% SWE-bench Verified at a Tenth of the Price
The open-model tier split three ways in July 2026. For frontend, Kimi K3 leads the Frontend Code Arena ahead of Fable 5 (weights due ~July 27). For value, GLM-5.2 is the top open model on the Artificial Analysis Intelligence Index, served on Morph at $1.10/$4.10. For raw SWE-bench Verified, DeepSeek-V4-Pro-Max is highest at 80.6% (vendor), tied with Gemini 3.1 Pro and self-hostable under MIT, with MiniMax M3 at 80.5% and Kimi K2.6 at 80.2%. Official API prices and self-host terms:
| Model | Benchmark | $/M in / out (official API) | Context / license |
|---|---|---|---|
| Kimi K3 (Moonshot, 2.8T MoE) | #1 Frontend Code Arena; 93.4% Verified (Vals AI) | $3.00 / $15.00 | 1M / weights due ~Jul 27 |
| DeepSeek-V4-Pro-Max (1.6T / 49B active) | 80.6% Verified (top downloadable open weights) | $0.435 / $0.87 (V4 Pro) | 1M / MIT |
| DeepSeek V4 Flash (284B / 13B active) | 79.0% (Flash-Max) | $0.14 / $0.28 | 1M / MIT |
| morph-dsv4flash (DeepSeek V4 Flash on Morph) | bf16 activations, codegen spec decoding + kernels | $0.139 / $0.278 | MIT weights, hosted |
| MiniMax M3 | 80.5% | $0.60 / $2.40 | ~1M / open weights |
| Kimi K2.6 (1T / 32B active) | 80.2% | $0.95 / $4.00 | 256K / open weights |
| Qwen3.6 Plus | 78.8% | $0.50 / $3.00 | 1M / Qwen |
| GLM-5.2 (744B / 40B active) | #1 open model, Artificial Analysis Intelligence Index; 62.1% Pro (vendor) | $1.40 / $4.40 | 1M / MIT |
| morph-glm52-744b (GLM-5.2 on Morph) | bf16 activations, codegen kernels | $1.1 / $4.1 | 1M, hosted |
The arithmetic that matters: MiniMax M3 produces output at $2.40/M against Opus 4.8's $25/M, a 10.4x gap, while trailing it by 8.1 points on SWE-bench Verified (80.5 vs 88.6). DeepSeek V4 Flash sets the absolute floor at $0.28/M output with a 1M-token context. The price gap is mostly model size, not provider margin, which is the core fact of how AI inference is priced. For teams with data-sovereignty requirements, DeepSeek V4's MIT license means the 80.6% model is self-hostable outright. Deeper coverage on the best open-source coding model page.
Open weights are identical everywhere, the serving stack is not. Most serverless providers quantize activations to fp8 to cut cost, which degrades output quality. Morph serves these models with 16-bit (bf16) activations and does not quantize them, so output matches the reference weights. That makes Morph the place to run open models when fidelity matters. For coding specifically, Morph adds speculative decoding tuned on code plus custom low-level inference kernels built for code generation. Four coding-relevant open models on the same OpenAI-compatible API: GLM-5.2 (morph-glm52-744b) at $1.1/$4.1, Qwen 3.5 397B (morph-qwen35-397b) at $0.5/$3.5, DeepSeek V4 Flash (morph-dsv4flash) at $0.139/$0.278, and MiniMax M3 (morph-minimax3-428b) at $0.6/$2.4. Full catalog on Morph models and pricing.
What People Actually Code With (r/LocalLLaMA, mid-2026)
Benchmark tables and the models people run every day have drifted apart. On r/LocalLLaMA, the open-weight coding conversation in mid-2026 is dominated by GLM-5.2 and Qwen 3.6, not by whichever model tops SWE-bench that week. The recurring theme: a model you can actually run, at a token budget you can afford, beats two extra benchmark points you pay frontier prices for.
The most-upvoted open-weight coding threads on the subreddit this cycle are about running GLM-5.2 locally, not about the leaderboard. "GLM-5.2 is a win for local AI" (1,200+ upvotes, 315 comments) and "GLM-5.2 on 5x Pro 6000s and a 5090, an expensive journey" (1,500+ upvotes) are the community reckoning with what it costs to self-host a 744B MoE well enough to code against. The Qwen side shows up as "Qwen3.6 35B-A3B (Q8_0, no KV quant) single prompt", a 35B active-3B model people run unquantized on a single box.
Two practitioner signals matter for a "best LLM for coding" pick. The GPU math is brutal: threads describe five RTX Pro 6000s plus a 5090 to serve GLM-5.2 at usable speed, which is why most teams that want the open model's output rent it hosted rather than build the rack. And the export-control drama around Claude Fable and gated GPT-5.6 ("US Govt to individually approve who gets GPT 5.6") is a real availability risk that pushed developers toward open weights they control. Running GLM-5.2 or DeepSeek V4 on a hosted OpenAI-compatible API like Morph is the middle path: the open model's output without the rack.
Best AI Model for Coding at $0
The best free path to real coding capability in July 2026 is DeepSeek V4's MIT-licensed open weights (self-host) or its near-free hosted API at $0.14/$0.28. Four no-credit-card options:
| Option | What you get | Limit |
|---|---|---|
| Codex CLI on ChatGPT Free | Codex CLI with GPT-5.5, $0/mo | Lowest usage limits; Plus ($20/mo) raises them |
| Z.AI free GLM tier | Free Flash-tier GLM models on the Z.AI API | Flash tier only; GLM-5.2 flagship is paid at $1.40/$4.40 |
| Qwen on Alibaba Model Studio | Free token allowance for new users across Qwen models | Time-limited trial |
| DeepSeek V4 open weights | MIT-licensed weights, self-host V4 Flash (284B/13B active) | Your GPU cost; hosted API is $0.14/$0.28 anyway |
The honest framing: free tiers are for evaluation and light use. DeepSeek V4 Flash's paid API at $0.14/M input ($0.0036/M on cache hits) and $0.28/M output is close enough to zero that most teams skip self-hosting unless data cannot leave their network.
Why Vendor Scores Run 20 Points Above Scale's Leaderboard
Anthropic reports Fable 5 at 80.0% on SWE-bench Pro. Scale's standardized leaderboard tops out at 61.50% (Muse Spark 1.1). Both numbers are real. The difference is the harness: Scale runs every model through identical scaffolding; vendors run their own tuned agent stacks. Scale has not yet run the newest frontier models (Fable 5, Opus 4.8, GPT-5.6, Kimi K3), so their Pro numbers below are vendor-reported. (Vendor SWE-bench Verified numbers are also self-reported; llm-stats lists 0 of 104 as independently verified, which is why the Vals AI single-harness numbers above are the cleaner Verified reference.)
Same Benchmark, Different Harness: SWE-bench Pro
Vendor-reported scaffolds vs Scale SEAL standardized scaffolding, July 2026.
The vendor-vs-standardized gap is roughly 20 points for the same model families. The harness is the variable.
The practical conclusion has not changed since 2025: the scaffold around the model accounts for more variance than swapping frontier models. Before paying a 2x token premium, fix retrieval, context management, and tool design. Subagent architecture and context engineering move scores more than model choice does.
A mid-tier model in a strong harness beats a frontier model in a weak one. Tools like WarpGrep (semantic codebase search for terminal agents, $0 for 100k requests) upgrade the harness for every model you route through it.
Per-Task Routing: Which Model for Which Job
The most cost-effective setups in July 2026 route by task, not by loyalty: send the hard 20% to Opus 4.8 and the cheap 80% to Haiku 4.5 or DeepSeek V4 Flash. Numbers-backed defaults:
| Task | Route to | Why (verified numbers) |
|---|---|---|
| Overnight refactor, 50+ files | Claude Opus 4.8 | 69.2% SWE-bench Pro (vendor), 1M context, no long-context surcharge |
| Hardest debugging / migration runs | GPT-5.6 Sol, Fable 5, or Opus 4.8 | 96.2% / 95.0% / 88.6% Verified (Vals AI); Opus 4.8 at the lowest price of the three |
| Quick edits, lint fixes, subagents | Claude Haiku 4.5 | $1/$5, ~$0.13 output per Pro point, $0.10/M cache hits |
| Terminal / Codex workflows | GPT-5.6 Sol | $5/$30, the model OpenAI now ships through Codex; 96.2% Verified (Vals AI) |
| Standardized-harness ceiling | gpt-5.4 | 59.10% Scale SEAL SWE-bench Pro at $2.50/$15 |
| High-volume batch / CI bots | DeepSeek V4 Flash or MiniMax M3 | $0.28/M and $2.40/M output, both ~1M context |
| Budget proprietary, long prompts | Gemini 3.1 Pro | 46.10% Scale Pro at $2/$12 (≤200K); input doubles to $4 above 200K |
| Data sovereignty / self-host | DeepSeek V4 (MIT) | 80.6% SWE-bench Verified (Pro Max), weights on Hugging Face |
| Codebase search for any agent | WarpGrep + any model | Model-agnostic retrieval; $0 for 100k requests |
Cost levers that apply across routes: Anthropic's Batch API is 50% off input and output, prompt-cache reads are 0.1x base input, and DeepSeek cache hits drop input to $0.0036/M. A routing setup that pins 80% of traffic to Haiku 4.5 or DeepSeek V4 Flash and reserves Opus 4.8 for the hard 20% typically beats any single-model subscription. Doing the split automatically needs a classifier: Morph's Router scores each prompt by difficulty and domain in ~180ms and returns the cheapest capable model, so a coding agent gets cheaper and faster at once (see the model lineup and pricing). Claude Code Router makes that per-request routing concrete inside the terminal agent, and Claude Code models covers harness-side defaults.
Frequently Asked Questions
What is the best AI model for coding in 2026?
It depends on the job. On Vals AI's independent SWE-bench Verified harness the top model in July 2026 is GPT-5.6 Sol at 96.2%, then Claude Fable 5 at 95.0% and Kimi K3 at 93.4%. Claude Opus 4.8 (claude-opus-4-8, 88.6% Verified, $5/$25, 1M context) is the everyday pick. For frontend work, Kimi K3 tops the Arena.ai Frontend Code Arena ahead of Fable 5. The best open-weight value is GLM-5.2 (top of the Artificial Analysis Intelligence Index, $1.10/$4.10 on Morph), with DeepSeek V4 (80.6% vendor Verified, MIT) alongside it. For most teams the best answer is not one model but a router that sends easy work to a cheap or open model and reserves a frontier model for hard edits.
What is the best LLM for coding?
Same question as best AI model for coding, and the ranking above answers both. On the independent Vals AI harness, GPT-5.6 Sol leads at 96.2% SWE-bench Verified, then Claude Fable 5 (95.0%) and Kimi K3 (93.4%); Claude Opus 4.8 (88.6%, $5/$25) is the everyday default. For frontend, Kimi K3 is #1 on the Arena.ai Frontend Code Arena. On Scale's standardized SWE-bench Pro leaderboard, Muse Spark 1.1 leads at 61.50%, ahead of gpt-5.4 (59.10%). Cost-adjusted, Claude Haiku 4.5 ($1/$5) is the cheapest per solved benchmark point at roughly $0.13 of output. For this same field ranked picks-by-job, see best LLM for coding.
Is there an open source LLM as good as Claude for coding?
On frontend work, one leads: Kimi K3 tops the Arena.ai Frontend Code Arena ahead of Claude Fable 5, and scores 93.4% SWE-bench Verified on Vals AI's harness (weights due on Hugging Face around July 27). On general SWE-bench Verified, the best downloadable open-weight model, DeepSeek-V4-Pro-Max, scores 80.6% against Opus 4.8's 88.6% and Fable 5's 95.0%. But the open models cost roughly a tenth as much per output token (GLM-5.2 $4.10 on Morph, DeepSeek V4 Flash $0.28, versus Opus at $25) and ship under permissive licenses you can self-host. For the 80% of coding work that is not the hardest debugging or migration, an open model at a tenth of the price is the better buy. Deeper comparison on the best open source LLM and best open source coding model pages.
Which Claude model is best for coding?
Claude Fable 5 (claude-fable-5, $10/$50) is the highest-scoring Claude at 95.0% SWE-bench Verified and is available again as of July 1. Claude Opus 4.8 (API ID claude-opus-4-8, $5/$25) is the everyday default at half the price: 88.6% SWE-bench Verified, 69.2% SWE-bench Pro on Anthropic's harness, 1M context with no long-context surcharge. Claude Sonnet 4.6 (claude-sonnet-4-6, $3/$15, 79.6% Verified) is the volume pick with a 1M context, and Claude Haiku 4.5 (claude-haiku-4-5, $1/$5) handles quick edits and subagents. Treat Sonnet 4, Opus 4, and Opus 4.1 as legacy and migrate off them.
What is the best Codex model for coding in 2026?
OpenAI's current top model is GPT-5.6 Sol, which reached general availability on July 9, 2026 and is available on the API and through Codex at $5/$30 per million tokens. It leads Vals AI's independent SWE-bench Verified harness at 96.2%. OpenAI points Codex at the general gpt-5.6 model rather than a dedicated -codex variant, and GPT-5.5 ($5/$30) remains available. On Scale's standardized SWE-bench Pro board the newest frontier models are not run yet, so the top entries there are Muse Spark 1.1 (61.50%) and gpt-5.4 (59.10%).
What are the SWE-bench Pro scores for coding models in 2026?
Scale SEAL public set (standardized scaffolding, July 2026): Muse Spark 1.1 61.50%, gpt-5.4 xHigh 59.10%, Muse Spark 55.00%, Opus 4.6 thinking 51.90%, Gemini 3.1 Pro thinking 46.10%, Opus 4.5 45.89%, Sonnet 4.5 43.60%, Gemini 3 Pro 43.30%, gpt-5 (High) 41.78%. The newest frontier models (GPT-5.6, Fable 5, Opus 4.8, Kimi K3) are not on Scale's board yet. Vendor-reported numbers run higher: Anthropic reports Fable 5 at 80.0% and Opus 4.8 at 69.2%, and Z.ai reports GLM-5.2 at 62.1%. All SWE-bench Verified numbers are vendor self-reported; llm-stats lists 0 as independently verified, so the Vals AI single-harness numbers are the cleaner Verified reference.
What is the best free LLM for coding?
Four real $0 paths in July 2026: Codex CLI is included with a ChatGPT Free sign-in (lowest usage limits); Z.AI offers free Flash-tier GLM models on its API; Alibaba Model Studio gives new users a free token allowance across Qwen models; and DeepSeek V4's MIT-licensed weights are self-hostable. If you have a GPU, the strongest genuinely free LLM for coding is a self-hosted open-weight model (DeepSeek V4 Flash or GLM-5.2). DeepSeek V4 Flash's paid API is near-free anyway at $0.14/M input, $0.28/M output.
What is the best open-source AI model for coding?
For frontend, Kimi K3 is strongest: #1 on the Arena.ai Frontend Code Arena, 93.4% SWE-bench Verified (Vals AI), weights due ~July 27. On general SWE-bench Verified, DeepSeek-V4-Pro-Max leads downloadable open weights at 80.6% (vendor), tied with Gemini 3.1 Pro, with MiniMax M3 at 80.5% and Kimi K2.6 at 80.2%. For value, GLM-5.2 (744B MoE / 40B active, MIT, 1M context) is the top open model on the Artificial Analysis Intelligence Index; it lists at $1.40/$4.40 on Z.ai and runs $1.10/$4.10 on Morph. DeepSeek V4 ships under MIT: V4 Pro (1.6T / 49B active) costs $0.435/$0.87, V4 Flash (284B/13B) costs $0.14/$0.28.
How much do the top coding models cost per million tokens?
Output price ladder, July 2026: DeepSeek V4 Flash $0.28, DeepSeek V4 Pro $0.87, GLM-5.2 $4.40 ($4.10 on Morph), Kimi K2.6 $4.00, Gemini 3.1 Pro $12, Kimi K3 $15 (Moonshot), Claude Sonnet 5 $10 (intro), Claude Opus 4.8 $25, GPT-5.5 and GPT-5.6 Sol $30, Claude Fable 5 $50. Inputs range from $0.14/M (DeepSeek V4 Flash) to $10/M (Fable 5). GLM-5.2 on Morph is $1.10 input / $4.10 output, below Z.ai's own $1.40/$4.40.
Why do vendor benchmark scores differ from Scale's leaderboard?
Scale runs every model through identical standardized scaffolding on SWE-bench Pro's 1,865 tasks across 41 repositories, scored Pass@1; vendors run their own tuned harnesses. The same model family scores 51.90% (Opus 4.6 on Scale) versus 69.2% (Opus 4.8 on Anthropic's harness). That roughly 20-point spread is the harness, which is why agent tooling moves results more than model swaps. The newest frontier models are not on Scale's board yet, so the independent Vals AI harness is the cleaner Verified reference for them.
Which AI model is most cost-effective for coding in 2026?
For raw per-token cost with a 1M context, DeepSeek V4 Flash at $0.14/$0.28 is the floor. For best capability-per-dollar among open models, GLM-5.2 at $1.10/$4.10 on Morph tops the Artificial Analysis Intelligence Index among open weights. Dividing output price by Scale SEAL SWE-bench Pro score: Claude Haiku 4.5 about $0.13 of output per point, gpt-5.4 $0.25, Gemini 3.1 Pro $0.26, Claude Opus 4.6 $0.48.
Sources
Primary sources behind the scores and prices on this page (updated July 24, 2026):
- Vals AI SWE-bench Verified (independent single-harness scores: GPT-5.6 Sol, Fable 5, Kimi K3, Opus 4.8)
- Arena.ai Frontend Code Arena (blind pairwise human votes; Kimi K3 #1)
- Scale SEAL SWE-bench Pro public leaderboard (standardized harness scores)
- llm-stats SWE-bench Verified tracker and SWE-bench Pro aggregate (vendor self-reported)
- SWE-bench Pro paper (1,865 tasks / 41 repos, benchmark definition)
- Anthropic Claude pricing (Fable 5 $10/$50, Opus 4.8 $5/$25)
- OpenAI GPT-5.6 (GA July 9) and OpenAI API pricing
- Moonshot Kimi K3 pricing and Artificial Analysis Kimi K3 (throughput / TTFT)
- Google Gemini API pricing
- DeepSeek API pricing and Z.ai GLM-5.2 docs
- Artificial Analysis: GLM-5.2 leads the open-weights Intelligence Index
- Artificial Analysis Coding Agent Index methodology (cost per task, token usage per task)
- r/LocalLLaMA (practitioner threads on GLM-5.2 and Qwen 3.6 for local coding, mid-2026)
Stop Debating Models. Start Searching Codebases.
WarpGrep adds semantic codebase search to any terminal agent. Works with GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro, DeepSeek V4, GLM-5.2, or any model. $0 for 100k requests, $1 per 1M on Pro. The harness matters more than the model.
