TL;DR
Last updated October 7, 2026.
DeepSeek V4 Flash is a 284B-parameter MoE with 13B active per token, a 1M-token context, and MIT weights. The GA checkpoint is DeepSeek-V4-Flash-0731. On September 10, 2026 DeepSeek retired it on its own API: the deepseek-v4-flash model name now routes to V4.1 Flash. To keep calling the 0731 weights, use a host that pins them, such as morph-dsv4flash at $0.141953125/M input and $0.399625/M output.
โoutput tokens for DeepSeek V4 Flash 0731 to run the Intelligence Index, against a 140M median for open-weight models of similar size.โ
What it is
284B total / 13B active MoE, 1M-token context, CSA + HCA sparse attention, text only. MIT weights on Hugging Face. Artificial Analysis Intelligence Index: 34, against an 18 median for its size. 220.8 output tokens/s on DeepSeek's API.
What changed in September
DeepSeek shipped V4.1 Flash (552B backbone, about 763B total with Engram memory and vision encoder, 8B/16B active, image input) on September 10 and pointed deepseek-v4-flash at it. OpenRouter's deepseek/deepseek-v4-flash slug is the April 0423 preview. The 0731 weights only answer where a host pins them.
DeepSeek V4 Flash Benchmarks at a Glance
The headline DeepSeek V4 Flash benchmark numbers, next to V4.1 Flash, which now answers the deepseek-v4-flash name on DeepSeek's API. DeepSeek reports the first four rows itself. Artificial Analysis runs its Intelligence Index independently. The per-source breakdown is in the benchmarks section.
| Benchmark | V4 Flash (0731) | V4.1 Flash | Source |
|---|---|---|---|
| GPQA Diamond (max effort) | 88.1 | 90.9 | DeepSeek, self-reported |
| Terminal-Bench 2.1 | 82.7 | 90.6 | DeepSeek, self-reported |
| DeepSWE | 54.4 | 74.2 (v1.1) | DeepSeek, self-reported |
| NL2Repo | 54.2 | 65.4 | DeepSeek, self-reported |
| AA Intelligence Index (max) | 34 (10th of 117) | 39 | Artificial Analysis, independent |
| AA output tokens to run the Index | 240M | 250M | Artificial Analysis, independent |
What Is DeepSeek V4 Flash?
DeepSeek V4 Flash is the smaller of the two DeepSeek V4 models. DeepSeek previewed it April 24, 2026 and released the GA checkpoint, DeepSeek-V4-Flash-0731, on July 31, 2026 under the MIT license. It has 284 billion total parameters with 13 billion active per token (V4 Pro is 1.6T / 49B), a 1M-token context window, and a recommended 384K max output for the high and max reasoning modes.
The 0731 checkpoint kept the architecture and rebuilt post-training for agents. On DeepSeek's model card, Terminal-Bench 2.1 went from 61.8 to 82.7 and DeepSWE from 7.3 to 54.4 against the preview. The reasoning_effort parameter takes three levels: low, high, and max. The legacy deepseek-chat and deepseek-reasoner aliases retired on July 24, 2026.
Flash was built for long context on small hardware. Per the Hugging Face DeepSeek V4 write-up, at 1M tokens Flash uses 10% of the single-token FLOPs and 7% of the KV cache of DeepSeek V3.2. For Pro the figures are 27% and 10%.
Morph serves the 0731 weights as morph-dsv4flash through an OpenAI-compatible API. See Morph Open Source Models and pricing. For the whole V4 family (Pro, Flash, vision), the hub is the DeepSeek V4 guide.
DeepSeek V4 Flash vs V4.1 Flash: Size, Benchmarks, and Which One Your API Call Hits
DeepSeek released V4.1 Flash on September 10, 2026. Its changelog says the previous-generation V4 Flash and V4 Flash Vision Exp "have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash." The new model name is deepseek-flash. Both legacy names bill at the V4.1 Flash price. If your code still sends deepseek-v4-flash to DeepSeek, it has been running a different model since that date, at different prices. DeepSeek's pricing page, checked October 7, 2026, still accepts the legacy names and serves them with V4.1 Flash. Code that sent images to deepseek-v4-flash-vision-exp keeps working, because V4.1 Flash reads images natively.
The same release note announced that deepseek-v4-pro would route to V4.1 Flash from September 14. DeepSeek reversed that in the changelog: it "decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." The pricing page still lists deepseek-v4-pro (V4-Pro-0813) at $0.66/M input and $1.98/M output off-peak, $1.32/M and $3.96/M at peak. DeepSeek has not given a V4.1 Pro date.
DeepSeek API changelog for Flash, July to September 2026
- July 24:
deepseek-chatanddeepseek-reasonerdiscontinued. - July 31: V4-Flash-0731 GA on the API, same 284B architecture, new post-training.
- August 13: V4-Pro GA (V4-Pro-0813). Native OpenAI Responses API support and three effort levels (low / high / max) for Pro and Flash.
- August 16: peak and off-peak pricing starts. Off-peak is half the peak rate.
- August 21:
deepseek-v4-flash-vision-explaunches with a free Files API. Images bill at up to 384 tokens each. - September 10: V4.1 Flash ships as
deepseek-flashwith native image input and lower prices. V4 Flash and V4 Flash Vision Exp are retired and their names route to V4.1 Flash. - After September 14: V4 Pro stays on the API at its own prices, reversing the planned reroute.
| Property | V4 Flash (0731) | V4.1 Flash |
|---|---|---|
| Released | July 31, 2026 | September 10, 2026 |
| Backbone parameters (card) | 284B | 552B |
| Extra modules | DSpark draft module | 196B Engram memory + vision encoder |
| Total parameters (HF safetensors) | ~304B | ~763B |
| Active parameters | 13B | 8B input / 16B output |
| Architecture | Decoder MoE, CSA + HCA | Causal encoder-decoder, 40 layers, CSA2 |
| Global KV cache | baseline | 890 bytes/token (about 1/4) |
| Image input | No | Yes |
| Reasoning effort | low / high / max | low / high / max on the API; integer 1-100 on the weights |
| Training tokens | >32T | 45T |
| DeepSWE (v1.1) | 54.4 | 74.2 |
| Terminal-Bench 2.1 | 82.7 | 90.6 |
| AA Intelligence Index | 34 | 39 |
| AA output tokens to run Index | 240M | 250M |
| Weights on disk | 162GB (Q8 GGUF) | ~510GB |
| DeepSeek API model name | retired (name routes to V4.1) | deepseek-flash |
| Morph model id | morph-dsv4flash | morph-dsv41flash |
| Morph input / output per 1M | $0.141953125 / $0.399625 | $0.12 / $0.8 |
V4.1 Flash wins on quality. DeepSeek's card claims it beats V4-Pro on most agentic benchmarks, and Artificial Analysis scores it 5 points above Flash-0731. It is not a free upgrade for everyone. It carries about 2.5x the parameters (about 763B vs 304B on Hugging Face, counting its 196B Engram memory), so it no longer fits the 128GB boxes that ran V4 Flash. vLLM's V4.1 tracking lists it as verified on H200, GB200, GB300, and MI350X, and a September 16 issue (vLLM #57144) notes there is no Ampere (A100/A800) path yet. It also emits slightly more output tokens per Index run (250M vs 240M), and it is a different model: evals and prompts tuned on 0731 need re-running.
A measured data point from Level1Techs, posted to Hacker News on September 18: on 4x RTX 6000 Blackwell Pro Max-Q with SGLang and an fp8 KV cache, V4.1 Flash decoded at 102.0 tok/s per request against 48.8 for V4 Flash 0731, and prefilled at 5,955 tok/s against 2,444. The V4.1 run used DSPARK speculative decoding and a 524,288-token context cap; the 0731 run had the full 1,048,576.
V4.1 Flash on Morph is morph-dsv41flash at $0.12/M input, $0.008/M cached, and $0.8/M output with a 1M context. The launch details are in DeepSeek V4.1 Flash on the Morph API. This page stays on V4 Flash 0731.
On DeepSeek's API, deepseek-v4-flash now means V4.1 Flash. On OpenRouter, deepseek/deepseek-v4-flash is labeled "DeepSeek V4 Flash 0423", the April preview, and deepseek/deepseek-v4-flash-0731 is the GA checkpoint (OpenRouter model list, September 22, 2026). On Morph, morph-dsv4flash is 0731 and morph-dsv41flash is V4.1. Log the model id your provider returns, not the one you sent.
DeepSeek V4 Flash Size: 284B Parameters, Weights, and 128GB Hardware
DeepSeek's release note and the V4 technical report (arXiv:2606.19348) give Flash as 284B total parameters with 13B activated. The Hugging Face page for the 0731 repo shows 304B params, because its counter sums every tensor in the safetensors files. Plan with 284B / 13B.
| Build | Size on disk | Memory needed |
|---|---|---|
| UD-IQ3_XXS GGUF (recommended) | 103GB | at least 110GB RAM |
| UD-Q8_K_XL GGUF (full original precision) | 162GB | at least 169GB RAM/VRAM |
| V4.1 Flash (for comparison) | ~510GB | 8x 80GB node (vLLM #57144) |
That is why "deepseek v4 flash 128gb" is a common search. In the Hacker News launch thread one user wrote "You can run it under 128gb, so a $3000 strix halo would do", and another posted a live demo on a 128GB MacBook. Unsloth recommends temperature 1.0 and top-p 1.0, or top-p 0.95 for agentic work. If you run llama.cpp with a quantized KV cache, use a build after July 7, 2026, when PR #25202 ("fix quantized kv-cache for dsv4") merged.
Architecture: CSA + HCA, FP4/FP8
DeepSeek V4 Flash pairs two interleaved sparse-attention mechanisms, Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), with a mixed FP4/FP8 weight format. That combination is what gets a 1M-token context to 10% of V3.2's per-token FLOPs. DeepSeek publishes the attention design and Flash's efficiency ratios. It does not publish Flash's layer split, so that is marked unconfirmed below.
| Property | V4 Flash | V4 Pro (for reference) |
|---|---|---|
| Total parameters | 284B | 1.6T |
| Active parameters / token | 13B | 49B |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Recommended max output | 384K tokens | 384K tokens |
| Attention | CSA + HCA | CSA + HCA |
| Single-token FLOPs vs V3.2 (1M ctx) | 10% | 27% |
| KV cache vs V3.2 (1M ctx) | 7% | 10% |
| Weights (preview card) | FP4 experts / FP8 | FP4 experts / FP8 |
| Training tokens | >32T | >32T |
| License | MIT | MIT |
CSA and HCA: how the KV cache shrinks
Per the Hugging Face DeepSeek V4 blog, the two attention paths compress the KV cache at different ratios and interleave across layers:
- CSA compresses KV entries 4x along the sequence dimension. A lightning indexer selects the top-k compressed blocks per query, and a sliding-window branch handles the most recent uncompressed tokens.
- HCA compresses KV entries 128x and drops sparse selection: every query attends densely over the short compressed sequence.
DeepSeek publishes the layer arrangement for Pro: a 61-layer stack where layers 0-1 are HCA and layers 2-60 alternate CSA and HCA. It does not publish Flash's split, so treat Flash's layer count and CSA/HCA arrangement as unconfirmed. V4.1 Flash replaced this stack with CSA2 and an FP4 main KV cache, which is where its 890 bytes per token comes from. At 1M tokens that is about 0.9GB of global KV cache.
Mixed-precision weights: FP4 experts
The preview Flash card describes FP4 MoE experts with FP8 for most other parameters. The 0731 repo lists BF16, F8_E4M3, I8, and F32 tensors, so check which files you pull before sizing memory. Stacking another aggressive quant on top (an IQ3 GGUF, an fp8 KV cache) is where the garbage-output reports in the serving section came from.
Confirmed first-party (model cards, release notes, arXiv:2606.19348, HF blog): 284B total / 13B active, 1M context, 384K recommended max output, CSA + HCA attention, 10% FLOPs / 7% KV cache vs V3.2, FP4/FP8 mixed weights in the preview, >32T training tokens, MIT. Not published for Flash: layer count, CSA/HCA layer split, and the lightning indexer's top-k.
DeepSeek V4 Flash Benchmarks: Self-Reported vs Independent
Most DeepSeek V4 Flash benchmark numbers come from DeepSeek. The preview model card reports SWE-bench Verified at 73.7 in non-thinking mode, 78.6 at high effort, and 79.0 at max. llm-stats repeats the 79.0 for DeepSeek-V4-Flash-Max (rank #17 of 116) and lists the 0423 preview at 78.6. For the 0731 GA checkpoint, DeepSeek reports DeepSWE 54.4, Terminal-Bench 2.1 82.7, and Toolathlon-Verified 70.3 on its own harness. The independent numbers are Vals.ai's SWE-bench Verified run, Code Arena votes, and the Artificial Analysis Intelligence Index.
| Benchmark | Flash score | Source | Type |
|---|---|---|---|
| SWE-bench Verified | 73.7 / 78.6 / 79.0 (non-think / high / max) | DeepSeek V4 Flash model card | Self-reported |
| LiveCodeBench | 55.2 / 88.4 / 91.6 | DeepSeek V4 Flash model card | Self-reported |
| GPQA Diamond | 71.2 / 87.4 / 88.1 | DeepSeek V4 Flash model card | Self-reported |
| Terminal-Bench 2.1 (0731) | 82.7 (preview: 61.8) | DeepSeek 0731 model card | Self-reported |
| DeepSWE (0731) | 54.4 (preview: 7.3) | DeepSeek 0731 model card | Self-reported |
| Toolathlon-Verified (0731) | 70.3 (preview: 49.7) | DeepSeek 0731 model card | Self-reported |
| NL2Repo (0731) | 54.2 (preview: 39.4) | DeepSeek 0731 model card | Self-reported |
| SWE-bench Verified (0731, bash-only) | 88.8 (11th of 88) | Vals.ai (archived September 1, 2026) | Independent |
| Code Arena / WebDev (flash-high) | 1580 (tied 23rd) | arena.ai, September 22, 2026 | Independent, human votes |
| AA Intelligence Index (0731, max) | 34 (median 18) | Artificial Analysis | Independent, standardized |
The GA card also compares 0731 with outside models on DeepSeek's harness: Terminal-Bench 2.1 82.7 vs 81.0 for GLM-5.2 and 85.0 for Opus 4.8, and DeepSWE 54.4 vs 46.2 and 58.0. Those are useful for ranking, not for absolute numbers. No independent harness has re-run 0731's Terminal-Bench or DeepSWE claims as of September 22, 2026.
Vals.ai did run SWE-bench Verified on 0731 with a bash-only harness: 88.8, 11th of 88 models, against 96.4 for V4 Pro-0813 in 2nd. Vals then archived the benchmark on September 1, 2026 because scores had saturated, so V4.1 Flash will not get a Vals number. Scale's SWE-bench Pro leaderboard has not added an entry since July 9, 2026 and has no DeepSeek V4 result.
The independent signal is Artificial Analysis. Flash-0731 at max effort scores 34 on the Intelligence Index, against an 18 median for open-weight models of similar size. V4 Pro-0813 scores 36 and V4.1 Flash 39. AA also clocks Flash-0731 at 220.8 output tokens per second on DeepSeek's API, against a 67.6 median. For the V4 Pro side of this story, see the DeepSeek V4 guide.
Cheap Tokens, Expensive Habits
Price per token does not decide whether Flash is cheap. Output tokens per task do. Artificial Analysis measured Flash-0731 emitting 240M output tokens to run its full Intelligence Index, against a 140M median for open-weight models of similar size. V4 Pro-0813 used 160M. The smaller model reasons longer to reach its answers.
| Model and rate | AA Index score | Output tokens | Output price / 1M | Index output cost |
|---|---|---|---|---|
| V4 Flash 0731 on Morph (morph-dsv4flash) | 34 | 240M | $0.399625 | $95.91 |
| V4.1 Flash, DeepSeek API off-peak | 39 | 250M | $0.60 | $150.00 |
| V4.1 Flash on Morph (morph-dsv41flash) | 39 | 250M | $0.8 | $200.00 |
| V4 Pro 0813, DeepSeek API off-peak | 36 | 160M | $1.98 | $316.80 |
This counts output tokens only and ignores input and cache, so read it as a ratio. Pro is terser, but it still costs more per Index run than either Flash at these rates. Within Flash, your effort setting is the main cost lever. 0731 has three levels (low / high / max) and AA ran max; pick the lowest level that clears your task. Then cache: on Morph, cached input on morph-dsv4flash is $0.0359375/M against $0.141953125/M uncached.
What makes a turn hit, how the cached rate compares across providers, and why routing a session across models resets the cache are covered on the prompt caching page.
Flash vs Pro: When Each Wins
Use Flash for high-volume and cost-sensitive work. Use Pro when a hard, long-horizon agent loop needs the extra active parameters. On Artificial Analysis the gap between Flash-0731 and Pro-0813 is 2 points (34 vs 36). Flash is also about 3x faster on DeepSeek's API: 220.8 vs 67.8 output tokens per second.
| Dimension | V4 Flash (0731) | V4 Pro (0813) | Who wins |
|---|---|---|---|
| Active params / token | 13B | 49B | Pro (capability) |
| AA Intelligence Index | 34 | 36 | Pro (+2) |
| AA output speed (DeepSeek API) | 220.8 tok/s | 67.8 tok/s | Flash (3.3x) |
| Output tokens to run the Index | 240M | 160M | Pro (terser) |
| SWE-bench Verified (max, self-reported) | 79.0 | 80.6 | Pro (+1.6) |
| SWE-bench Verified (Vals.ai, bash-only) | 88.8 | 96.4 | Pro (+7.6) |
| Self-host on a 128GB box | Yes (quantized) | No (multi-node) | Flash |
| On DeepSeek's API today | Retired (routes to V4.1) | Served, $1.98/M output off-peak, $3.96/M peak | Pro |
The routing pattern most teams land on: Flash for bulk extraction, classification, first-draft edits, and background agent tasks; Pro for the hardest planning and multi-file refactors. Both use the same API shape and a 1M context, so switching is a model-string change.
Serving Footguns (vLLM / SGLang)
Flash gets its efficiency from low-precision attention and weights, and that is where self-hosted serving breaks. The serving flags are --tool-call-parser deepseek_v4 and --reasoning-parser deepseek_v4. Status of the filed issues as of September 22, 2026:
- FP8 crashes under DP + EP (vLLM #43648), still open. DeepSeek-V4-Flash-FP8 on vLLM v0.20.1 with data-parallel + expert-parallel on H200 crashes after processing part of a benchmark run.
- DSML tool parser mishandles wrapped/reserved args (vLLM #41240), closed May 6, 2026. Wrapper params named
argumentsorinputwere unwrapped when they were real schema fields. Upgrade past the fix if your tools use those names. - FP8 + pipeline-parallel garbage output (SGLang #25662), closed May 19, 2026. DeepSeek-V4-Flash-FP8 with
--pp-size 8on 8x H20 emitted garbage CJK tokens. Older SGLang builds still have it. - llama.cpp quantized KV cache (PR #25202), merged July 7, 2026. Builds before that date mishandle a quantized KV cache on DeepSeek V4.
- V4.1 Flash is a separate code path. vLLM registers it as
DeepseekV41ForCausalLM. One September 18 report (vLLM #57469) hit a host-memory OOM loading its ~510GB of weights on 4x H20 with 587GiB of node RAM. Budget host RAM, not just HBM, if you move up.
Validate your exact quant-plus-parallelism combination on a small eval before production. Keep the reasoning content in history: DeepSeek's encoding scripts retain reasoning_content that a naive Jinja template would strip, which breaks multi-turn tool calling.
Pricing and How to Run It on Morph
DeepSeek no longer sells V4 Flash 0731 by name. Its pricing page lists deepseek-flash (V4.1 Flash) at $0.15/M input (cache miss), $0.003/M (cache hit), and $0.60/M output off-peak, doubling to $0.30 / $0.006 / $1.20 during peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays. The new rates took effect at 04:00 UTC on September 10. The legacy deepseek-v4-flash name bills at those rates. Artificial Analysis still lists Flash-0731 at $0.44/M input and $1.32/M output, the peak rate before September 10.
OpenRouter lists deepseek/deepseek-v4-flash-0731 from $0.04/M input and $0.64/M output (lowest-priced host, September 22, 2026).
Morph serves morph-dsv4flash on the 0731 checkpoint at 16-bit (bf16) activations with no fp8 activation quantization, at $0.141953125/M input and $0.399625/M output, flat at every hour, with a 1M context. For coding agents, Morph adds codegen-tuned speculative decoding and custom inference kernels. Call it through the OpenAI-compatible API:
from openai import OpenAI
client = OpenAI(
base_url="https://api.morphllm.com/v1",
api_key="YOUR_MORPH_API_KEY",
)
resp = client.chat.completions.create(
model="morph-dsv4flash",
messages=[
{"role": "user", "content": "Refactor this function to be async."},
],
)
print(resp.choices[0].message.content)See Morph Open Source Models for the full lineup and pricing for every model's rates. For the broader DeepSeek V4 family and Claude Code setup, see the DeepSeek V4 guide and DeepSeek API guide.
DeepSeek V4 Flash API: Pricing and How to Call It
The DeepSeek V4 Flash API is OpenAI-compatible, so calling it is a base-URL and model-id change. The model id decides which checkpoint you get. On DeepSeek, deepseek-v4-flash now answers with V4.1 Flash. On Morph, morph-dsv4flash answers with the 0731 weights at $0.141953125/M input and $0.399625/M output with the full 1M context.
Call the DeepSeek V4 Flash API with the OpenAI SDK
Point any OpenAI-SDK client at https://api.morphllm.com/v1, pass your key as a Bearer token, and set the model to morph-dsv4flash. The raw HTTP call is a POST to https://api.morphllm.com/v1/chat/completions with an Authorization: Bearer YOUR_MORPH_API_KEY header. Flash is verbose, so set the reasoning effort to the lowest level that clears your task, and keep the model's reasoning_content in history so multi-turn tool calling does not break.
curl https://api.morphllm.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_MORPH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "morph-dsv4flash",
"messages": [{"role": "user", "content": "Summarize this diff."}]
}'DeepSeek V4 Flash API pricing compared
| Provider | Model id | Checkpoint | Input / 1M | Cache-hit input | Output / 1M |
|---|---|---|---|---|---|
| Morph | morph-dsv4flash | V4 Flash 0731 | $0.141953125 (flat) | $0.0359375 | $0.399625 (flat) |
| Morph | morph-dsv41flash | V4.1 Flash | $0.12 (flat) | $0.008 | $0.8 (flat) |
| OpenRouter (lowest host) | deepseek/deepseek-v4-flash-0731 | V4 Flash 0731 | $0.04 | varies | $0.64 |
| DeepSeek (off-peak) | deepseek-flash / deepseek-v4-flash | V4.1 Flash | $0.15 | $0.003 | $0.60 |
| DeepSeek (peak) | deepseek-flash / deepseek-v4-flash | V4.1 Flash | $0.30 | $0.006 | $1.20 |
Sources: DeepSeek's Models and Pricing page and change log; OpenRouter's public model list. Morph rates are the canonical pricing.
Getting a DeepSeek V4 API key on Morph
Sign up at morphllm.com, create a key in the dashboard, and send it as Authorization: Bearer YOUR_MORPH_API_KEY. One key covers the whole model lineup, including both Flash checkpoints. For the DeepSeek API walkthrough (setup, cache pricing, Claude Code), see the DeepSeek API guide.
Flash vs GLM 5.2 and GLM-5.3-Flash, MiniMax M3, Qwen 3.5 and Qwen3.8
Among open-weight coding models, V4 Flash 0731 is the one that fits a 1M context on a single 128GB box. Its successor V4.1 Flash scores higher but needs a full GPU node. DeepSeek's own 0731 card has GLM-5.2 at 81.0 on Terminal-Bench 2.1 against Flash-0731's 82.7, and 46.2 on DeepSWE against 54.4. Both GLM 5.2 and Qwen 3.5 now have successors. Z.ai shipped GLM-5.3 on August 18 and GLM-5.3-Flash on August 26, 2026. Alibaba's current line is Qwen3.8.
| Model | Size (total / active) | License | Vendor API input / output per 1M | AA Index |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 284B / 13B | MIT | retired on DeepSeek; routes to V4.1 | 34 |
| DeepSeek V4.1 Flash | ~763B (552B backbone) / 8B-16B | MIT | $0.15 / $0.60 off-peak | 39 |
| GLM-5.3-Flash | 320B / 18B | MIT | $0.15 / $0.50 | 42 |
| MiniMax-M3 | ~428B / ~23B | minimax-community | $0.30 / $1.20 (up to 512K input) | 29 |
| Qwen3.8-Flash (hosted; Flash-Next weights) | ~180B / 6B | qwen-community-1.0 | $0.15 / $0.47 | 40 |
GLM-5.3-Flash is the closest match: MIT weights, 1M context, image and video input, and the highest Index score of the group at the same $0.15 input price as V4.1 Flash. MiniMax M3 also reads images and video but trails on the Index. Qwen 3.5 ships under Apache 2.0, and so does the dense Qwen3.8-27B. The MoE Qwen3.8 weights use custom licenses.
Deep dives: GLM 5.2, GLM-5.3-Flash, MiniMax M3, and Qwen 3.5. For a ranked list, see the best open-source coding models in 2026.
Pros and Cons
- Low per-token price for the 0731 weights ($0.141953125/$0.399625 per M on Morph; from $0.04/$0.64 on OpenRouter)
- Fits on a 128GB machine quantized (103GB IQ3 GGUF); V4.1 Flash does not
- MIT license; preview and 0731 weights on Hugging Face
- AA Intelligence Index 34 against an 18 median for its size, 2 points behind V4 Pro-0813
- 220.8 output tokens/s on DeepSeek's API, against a 67.6 median
- Pinned checkpoint: evals and prompts tuned on 0731 stay valid on hosts that keep serving it
- Retired on DeepSeek's API: deepseek-v4-flash routes to V4.1 Flash since September 10, 2026
- Verbose: 240M output tokens on the AA Index vs a 140M median
- 5 points behind V4.1 Flash on the AA Index (34 vs 39)
- 0731 agentic gains (Terminal-Bench 82.7, DeepSWE 54.4) are self-reported
- Text only; V4.1 Flash reads images
- FP8 under vLLM DP + EP still crashes on H200 (vLLM #43648, open)
When to Use It, When Not
Use DeepSeek V4 Flash 0731 when you have evals, prompts, or fine-tunes built on it and need the same weights answering tomorrow. Use it when you self-host on one 128GB machine, or when output price matters more than 5 Index points: bulk extraction, classification, first-draft edits, and background agent tasks.
Move to V4.1 Flash when you want the higher scores or image input and can re-run your evals. Read the V4.1 Flash launch post first. Reach for V4 Pro on long-horizon agent loops where terseness matters. The self-hosting numbers for Pro are on the DeepSeek V4 Pro serving benchmarks section. On any task where Flash's verbosity blows your token budget, cap it at the low or high reasoning effort level and lean on caching.
Frequently Asked Questions
What is DeepSeek V4 Flash?
The smaller DeepSeek V4 model, previewed April 24, 2026 and GA on July 31, 2026 as DeepSeek-V4-Flash-0731, under MIT. 284B total / 13B active, 1M-token context, 384K recommended max output. DeepSeek retired it on its own API on September 10, 2026 in favor of V4.1 Flash. The weights stay on Hugging Face, and third-party hosts still serve them.
Does deepseek-v4-flash still call V4 Flash on DeepSeek's API?
No. Since September 10, 2026 the name is temporarily routed to V4.1 Flash and billed at the Flash price, and so is deepseek-v4-flash-vision-exp. The new name is deepseek-flash. To keep the 0731 weights, call deepseek/deepseek-v4-flash-0731 on OpenRouter or morph-dsv4flash on Morph.
What is the difference between DeepSeek V4 Flash and V4.1 Flash?
V4 Flash is 284B / 13B active, text only. V4.1 Flash is a causal encoder-decoder with a 552B backbone plus 196B of Engram memory and a vision encoder, about 763B total on Hugging Face. It activates 8B for input and 16B for output, reads images, and uses about a quarter of the KV cache. DeepSWE v1.1: 54.4 vs 74.2 (DeepSeek). AA Intelligence Index: 34 vs 39.
How big is DeepSeek V4 Flash?
284B total parameters, 13B active (Hugging Face's counter shows 304B for the 0731 repo). Unsloth's GGUFs are 103GB (UD-IQ3_XXS, at least 110GB RAM) to 162GB (UD-Q8_K_XL). V4.1 Flash has a 552B backbone and about 763B parameters in total, about 510GB on disk.
How is Flash different from V4 Pro?
Pro is 1.6T / 49B; Flash is 284B / 13B. AA scores Pro-0813 at 36 and Flash-0731 at 34, and measured Flash at 220.8 output tok/s vs Pro's 67.8. Pro used 160M output tokens to run the Index, Flash 240M. DeepSeek still serves Pro at $0.66/M input and $1.98/M output off-peak, or $1.32/M and $3.96/M at peak. On Vals.ai's bash-only SWE-bench Verified run, Pro-0813 scored 96.4 and Flash-0731 88.8.
What is DeepSeek V4 Flash's SWE-bench score?
73.7 (non-thinking), 78.6 (high), and 79.0 (max) on SWE-bench Verified, all from DeepSeek's model card. llm-stats lists Flash-Max at 79.0, rank #17. The independent number is Vals.ai: 88.8 for 0731 on a bash-only run, 11th of 88, before Vals archived the benchmark on September 1, 2026. For 0731, DeepSeek also reports DeepSWE 54.4 and Terminal-Bench 2.1 82.7 on its own harness. No independent tracker has re-run those two.
Does cheap Flash pricing mean cheap tasks?
Not automatically. AA measured Flash-0731 emitting 240M output tokens to run its Index against a 140M median for its size. Cap the reasoning effort level and cache to control it.
Can DeepSeek V4 Flash run locally?
Yes. Unsloth's UD-IQ3_XXS quant is 103GB and needs at least 110GB RAM. Hacker News users report running it under 128GB, including on a 128GB MacBook. Use a llama.cpp build after July 7, 2026 (PR #25202) if you quantize the KV cache.
Is there a free DeepSeek V4 Flash API?
The weights are free under MIT. The hosted APIs are paid. As of October 7, 2026, OpenRouter still resolves a deepseek/deepseek-v4-flash:free slug (the 0423 preview), but it lists zero live endpoints, so it serves no requests. OpenRouter's paid slugs are deepseek/deepseek-v4-flash (0423), deepseek/deepseek-v4-flash-0731, and deepseek/deepseek-v4.1-flash. DeepSeek bills deepseek-v4-flash at V4.1 Flash rates ($0.15/M input, $0.60/M output off-peak).
What serving flags does Flash need on vLLM/SGLang?
--tool-call-parser deepseek_v4 and --reasoning-parser deepseek_v4, with an FP8 KV cache via --kv-cache-dtype fp8. vLLM #41240 and SGLang #25662 are fixed; the FP8 DP + EP crash (vLLM #43648) is still open.
Is DeepSeek V4 Flash open source?
Yes. The preview and 0731 weights are on Hugging Face under MIT. DeepSeek ships Python encoding scripts instead of a Jinja chat template. The technical report is arXiv:2606.19348.
When were deepseek-chat and deepseek-reasoner retired?
July 24, 2026, 15:59 UTC. Since September 10, 2026, deepseek-v4-flash itself routes to V4.1 Flash. Current DeepSeek model names are deepseek-flash and deepseek-v4-pro.
The efficiency default for coding agents
Use it when cost per task and interactive speed matter more than winning the hardest frontier task.
Best fits
- โ Interactive coding agents
- โ High volume routine work
- โ Long sessions with strong prompt reuse
Escalate or test carefully
- โข Kernel engineering
- โข Real world robotics
- โข Temporal reasoning
- โข Spatial reasoning
Cache the full agent session
Long coding sessions reuse system prompts, repository context, tool output, and prior turns. A useful production stack tiers that cache across GPU memory, CPU memory, and NVMe. GPU only cache sizing misses much of the cost per task opportunity.
Start with B300 when user speed matters, or B200 for aggregate throughput
Choose hardware around required speed per active user, then measure total capacity inside that latency target. Large batch throughput alone can hide a slow agent experience.
Related Articles
- DeepSeek V4: Full Guide (Pro + Flash, Architecture, Cost Math)
- DeepSeek V4.1 Flash on the Morph API
- DeepSeek API: Setup, Pricing, and Claude Code
- GLM 5.2: 753B Open-Weight Coding Model
- GLM-5.3-Flash
- MiniMax M3: 428B Multimodal Coding Agent
- Qwen 3.5: 397B MoE, Hybrid Attention
- Best Open-Source Coding Models in 2026
Plan a DeepSeek V4 Flash endpoint
Turn your team size and agent workload into a capacity estimate. Then validate the recommendation with your own traces.
An exact Morph capacity measurement is required before recommending a dedicated plan.
B200 is the compatible public platform. Dedicated capacity is invoiced monthly at the beginning of the month. Tokens are not billed separately.
Difference from serverless: $13,953 more per month.
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.
Run DeepSeek V4 Flash on Morph's OpenAI-Compatible API
morph-dsv4flash (V4 Flash 0731) at $0.141953125/M input and $0.399625/M output with a 1M context, or morph-dsv41flash (V4.1 Flash) at $0.12/$0.8. Pair either with WarpGrep so the context fills with the right code. WarpGrep runs $0.8 per 100K tokens.
Sources
- DeepSeek API Docs: Change Log (V4.1 Flash release, deepseek-v4-flash routing, V4 Pro continuation)
- DeepSeek API Docs: DeepSeek-V4.1-Flash release note (September 10, 2026)
- DeepSeek API Docs: Models and Pricing (deepseek-flash, deepseek-v4-pro, peak hours)
- DeepSeek API Docs: V4 Preview Release (April 24, 2026)
- Hugging Face: deepseek-ai/DeepSeek-V4-Flash-0731 model card (GA benchmarks, MIT)
- Hugging Face: deepseek-ai/DeepSeek-V4-Flash model card (params, FP4/FP8, SWE-bench by mode)
- Hugging Face: deepseek-ai/DeepSeek-V4.1-Flash model card (552B backbone, 196B Engram, ~763B total, CED, KV cache, benchmarks)
- Hugging Face blog: DeepSeek-V4, a million-token context (CSA / HCA, Flash FLOPs and KV ratios)
- arXiv:2606.19348: DeepSeek-V4, Towards Highly Efficient Million-Token Context Intelligence
- Artificial Analysis: DeepSeek V4 Flash 0731 (Intelligence Index, output tokens, speed; checked September 22, 2026)
- Artificial Analysis: DeepSeek V4.1 Flash (Intelligence Index 39, 250M output tokens)
- Artificial Analysis: DeepSeek V4 Pro 0813 (Intelligence Index 36, 160M output tokens)
- Vals.ai: SWE-bench Verified, bash-only (V4 Flash 0731 88.8, V4 Pro 0813 96.4; archived September 1, 2026)
- arena.ai: Code Arena / WebDev leaderboard (September 22, 2026)
- Z.ai: pricing (GLM-5.3-Flash)
- MiniMax: pay-as-you-go pricing (MiniMax-M3)
- Qwen Cloud: qwen3.8-flash pricing
- llm-stats: SWE-bench Verified leaderboard (Flash-Max 79.0, Pro-Max 80.6)
- OpenRouter: DeepSeek V4 Flash 0731 (per-host pricing)
- Level1Techs: V4.1 Flash vs V4 Flash on 4x RTX 6000 Blackwell Pro (SGLang throughput)
- vLLM #57144: SM8x (A100/A800) support for DeepSeek-V4.1-Flash
- vLLM #57469: DeepSeek-V4.1-Flash host-memory OOM on 4x H20
- vLLM #41240: DeepSeek V4 DSML tool parser mishandles wrapped and reserved arguments (closed)
- vLLM #43648: DeepSeek-V4-Flash-FP8 crashes under data-parallel + expert-parallel (open)
- SGLang #25662: Precision issues in DeepSeek V4 (closed)
- llama.cpp #25202: fix quantized kv-cache for dsv4 (merged July 7, 2026)
- Unsloth: DeepSeek V4 local serving notes (GGUF sizes, RAM, sampling)
- Hacker News: DeepSeek V4 discussion (running Flash under 128GB)