DeepSeek V4 Flash: 284B MoE, 1M Context, Benchmarks, Pricing

DeepSeek V4 Flash (0731) is a 284B/13B-active MIT model with a 1M context. Since September 10, 2026, DeepSeek's API routes deepseek-v4-flash to V4.1 Flash. Specs, size, benchmarks, pricing, and how to keep calling the 0731 weights (morph-dsv4flash at $0.141953125/$0.399625 per M).

July 12, 2026 ยท 1 min read

TL;DR

Last updated October 7, 2026.

DeepSeek V4 Flash is a 284B-parameter MoE with 13B active per token, a 1M-token context, and MIT weights. The GA checkpoint is DeepSeek-V4-Flash-0731. On September 10, 2026 DeepSeek retired it on its own API: the deepseek-v4-flash model name now routes to V4.1 Flash. To keep calling the 0731 weights, use a host that pins them, such as morph-dsv4flash at $0.141953125/M input and $0.399625/M output.

240M
โ€œoutput tokens for DeepSeek V4 Flash 0731 to run the Intelligence Index, against a 140M median for open-weight models of similar size.โ€
Artificial Analysis, DeepSeek V4 Flash 0731 model page, checked September 22, 2026

What it is

284B total / 13B active MoE, 1M-token context, CSA + HCA sparse attention, text only. MIT weights on Hugging Face. Artificial Analysis Intelligence Index: 34, against an 18 median for its size. 220.8 output tokens/s on DeepSeek's API.

What changed in September

DeepSeek shipped V4.1 Flash (552B backbone, about 763B total with Engram memory and vision encoder, 8B/16B active, image input) on September 10 and pointed deepseek-v4-flash at it. OpenRouter's deepseek/deepseek-v4-flash slug is the April 0423 preview. The 0731 weights only answer where a host pins them.

DeepSeek V4 Flash Benchmarks at a Glance

The headline DeepSeek V4 Flash benchmark numbers, next to V4.1 Flash, which now answers the deepseek-v4-flash name on DeepSeek's API. DeepSeek reports the first four rows itself. Artificial Analysis runs its Intelligence Index independently. The per-source breakdown is in the benchmarks section.

DeepSeek V4 Flash 0731 vs V4.1 Flash benchmarks (DeepSeek model cards and changelog, Artificial Analysis, checked October 7, 2026)
BenchmarkV4 Flash (0731)V4.1 FlashSource
GPQA Diamond (max effort)88.190.9DeepSeek, self-reported
Terminal-Bench 2.182.790.6DeepSeek, self-reported
DeepSWE54.474.2 (v1.1)DeepSeek, self-reported
NL2Repo54.265.4DeepSeek, self-reported
AA Intelligence Index (max)34 (10th of 117)39Artificial Analysis, independent
AA output tokens to run the Index240M250MArtificial Analysis, independent

What Is DeepSeek V4 Flash?

DeepSeek V4 Flash is the smaller of the two DeepSeek V4 models. DeepSeek previewed it April 24, 2026 and released the GA checkpoint, DeepSeek-V4-Flash-0731, on July 31, 2026 under the MIT license. It has 284 billion total parameters with 13 billion active per token (V4 Pro is 1.6T / 49B), a 1M-token context window, and a recommended 384K max output for the high and max reasoning modes.

The 0731 checkpoint kept the architecture and rebuilt post-training for agents. On DeepSeek's model card, Terminal-Bench 2.1 went from 61.8 to 82.7 and DeepSWE from 7.3 to 54.4 against the preview. The reasoning_effort parameter takes three levels: low, high, and max. The legacy deepseek-chat and deepseek-reasoner aliases retired on July 24, 2026.

Flash was built for long context on small hardware. Per the Hugging Face DeepSeek V4 write-up, at 1M tokens Flash uses 10% of the single-token FLOPs and 7% of the KV cache of DeepSeek V3.2. For Pro the figures are 27% and 10%.

$0.141953125 / $0.399625
morph-dsv4flash (V4 Flash 0731) input / output per 1M tokens, 1M context

Morph serves the 0731 weights as morph-dsv4flash through an OpenAI-compatible API. See Morph Open Source Models and pricing. For the whole V4 family (Pro, Flash, vision), the hub is the DeepSeek V4 guide.

DeepSeek V4 Flash vs V4.1 Flash: Size, Benchmarks, and Which One Your API Call Hits

DeepSeek released V4.1 Flash on September 10, 2026. Its changelog says the previous-generation V4 Flash and V4 Flash Vision Exp "have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash." The new model name is deepseek-flash. Both legacy names bill at the V4.1 Flash price. If your code still sends deepseek-v4-flash to DeepSeek, it has been running a different model since that date, at different prices. DeepSeek's pricing page, checked October 7, 2026, still accepts the legacy names and serves them with V4.1 Flash. Code that sent images to deepseek-v4-flash-vision-exp keeps working, because V4.1 Flash reads images natively.

The same release note announced that deepseek-v4-pro would route to V4.1 Flash from September 14. DeepSeek reversed that in the changelog: it "decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." The pricing page still lists deepseek-v4-pro (V4-Pro-0813) at $0.66/M input and $1.98/M output off-peak, $1.32/M and $3.96/M at peak. DeepSeek has not given a V4.1 Pro date.

DeepSeek API changelog for Flash, July to September 2026

  • July 24: deepseek-chat and deepseek-reasoner discontinued.
  • July 31: V4-Flash-0731 GA on the API, same 284B architecture, new post-training.
  • August 13: V4-Pro GA (V4-Pro-0813). Native OpenAI Responses API support and three effort levels (low / high / max) for Pro and Flash.
  • August 16: peak and off-peak pricing starts. Off-peak is half the peak rate.
  • August 21: deepseek-v4-flash-vision-exp launches with a free Files API. Images bill at up to 384 tokens each.
  • September 10: V4.1 Flash ships as deepseek-flash with native image input and lower prices. V4 Flash and V4 Flash Vision Exp are retired and their names route to V4.1 Flash.
  • After September 14: V4 Pro stays on the API at its own prices, reversing the planned reroute.
DeepSeek V4 Flash 0731 vs V4.1 Flash (DeepSeek model cards, Artificial Analysis, checked September 22, 2026)
PropertyV4 Flash (0731)V4.1 Flash
ReleasedJuly 31, 2026September 10, 2026
Backbone parameters (card)284B552B
Extra modulesDSpark draft module196B Engram memory + vision encoder
Total parameters (HF safetensors)~304B~763B
Active parameters13B8B input / 16B output
ArchitectureDecoder MoE, CSA + HCACausal encoder-decoder, 40 layers, CSA2
Global KV cachebaseline890 bytes/token (about 1/4)
Image inputNoYes
Reasoning effortlow / high / maxlow / high / max on the API; integer 1-100 on the weights
Training tokens>32T45T
DeepSWE (v1.1)54.474.2
Terminal-Bench 2.182.790.6
AA Intelligence Index3439
AA output tokens to run Index240M250M
Weights on disk162GB (Q8 GGUF)~510GB
DeepSeek API model nameretired (name routes to V4.1)deepseek-flash
Morph model idmorph-dsv4flashmorph-dsv41flash
Morph input / output per 1M$0.141953125 / $0.399625$0.12 / $0.8

V4.1 Flash wins on quality. DeepSeek's card claims it beats V4-Pro on most agentic benchmarks, and Artificial Analysis scores it 5 points above Flash-0731. It is not a free upgrade for everyone. It carries about 2.5x the parameters (about 763B vs 304B on Hugging Face, counting its 196B Engram memory), so it no longer fits the 128GB boxes that ran V4 Flash. vLLM's V4.1 tracking lists it as verified on H200, GB200, GB300, and MI350X, and a September 16 issue (vLLM #57144) notes there is no Ampere (A100/A800) path yet. It also emits slightly more output tokens per Index run (250M vs 240M), and it is a different model: evals and prompts tuned on 0731 need re-running.

A measured data point from Level1Techs, posted to Hacker News on September 18: on 4x RTX 6000 Blackwell Pro Max-Q with SGLang and an fp8 KV cache, V4.1 Flash decoded at 102.0 tok/s per request against 48.8 for V4 Flash 0731, and prefilled at 5,955 tok/s against 2,444. The V4.1 run used DSPARK speculative decoding and a 524,288-token context cap; the 0731 run had the full 1,048,576.

V4.1 Flash on Morph is morph-dsv41flash at $0.12/M input, $0.008/M cached, and $0.8/M output with a 1M context. The launch details are in DeepSeek V4.1 Flash on the Morph API. This page stays on V4 Flash 0731.

Gotcha: three model ids, three different checkpoints

On DeepSeek's API, deepseek-v4-flash now means V4.1 Flash. On OpenRouter, deepseek/deepseek-v4-flash is labeled "DeepSeek V4 Flash 0423", the April preview, and deepseek/deepseek-v4-flash-0731 is the GA checkpoint (OpenRouter model list, September 22, 2026). On Morph, morph-dsv4flash is 0731 and morph-dsv41flash is V4.1. Log the model id your provider returns, not the one you sent.

DeepSeek V4 Flash Size: 284B Parameters, Weights, and 128GB Hardware

DeepSeek's release note and the V4 technical report (arXiv:2606.19348) give Flash as 284B total parameters with 13B activated. The Hugging Face page for the 0731 repo shows 304B params, because its counter sums every tensor in the safetensors files. Plan with 284B / 13B.

DeepSeek V4 Flash weight sizes and memory (Unsloth docs, checked September 22, 2026)
BuildSize on diskMemory needed
UD-IQ3_XXS GGUF (recommended)103GBat least 110GB RAM
UD-Q8_K_XL GGUF (full original precision)162GBat least 169GB RAM/VRAM
V4.1 Flash (for comparison)~510GB8x 80GB node (vLLM #57144)

That is why "deepseek v4 flash 128gb" is a common search. In the Hacker News launch thread one user wrote "You can run it under 128gb, so a $3000 strix halo would do", and another posted a live demo on a 128GB MacBook. Unsloth recommends temperature 1.0 and top-p 1.0, or top-p 0.95 for agentic work. If you run llama.cpp with a quantized KV cache, use a build after July 7, 2026, when PR #25202 ("fix quantized kv-cache for dsv4") merged.

Architecture: CSA + HCA, FP4/FP8

DeepSeek V4 Flash pairs two interleaved sparse-attention mechanisms, Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), with a mixed FP4/FP8 weight format. That combination is what gets a 1M-token context to 10% of V3.2's per-token FLOPs. DeepSeek publishes the attention design and Flash's efficiency ratios. It does not publish Flash's layer split, so that is marked unconfirmed below.

DeepSeek V4 Flash architecture (HF model card, HF blog, arXiv:2606.19348)
PropertyV4 FlashV4 Pro (for reference)
Total parameters284B1.6T
Active parameters / token13B49B
Context window1,000,000 tokens1,000,000 tokens
Recommended max output384K tokens384K tokens
AttentionCSA + HCACSA + HCA
Single-token FLOPs vs V3.2 (1M ctx)10%27%
KV cache vs V3.2 (1M ctx)7%10%
Weights (preview card)FP4 experts / FP8FP4 experts / FP8
Training tokens>32T>32T
LicenseMITMIT

CSA and HCA: how the KV cache shrinks

Per the Hugging Face DeepSeek V4 blog, the two attention paths compress the KV cache at different ratios and interleave across layers:

  • CSA compresses KV entries 4x along the sequence dimension. A lightning indexer selects the top-k compressed blocks per query, and a sliding-window branch handles the most recent uncompressed tokens.
  • HCA compresses KV entries 128x and drops sparse selection: every query attends densely over the short compressed sequence.

DeepSeek publishes the layer arrangement for Pro: a 61-layer stack where layers 0-1 are HCA and layers 2-60 alternate CSA and HCA. It does not publish Flash's split, so treat Flash's layer count and CSA/HCA arrangement as unconfirmed. V4.1 Flash replaced this stack with CSA2 and an FP4 main KV cache, which is where its 890 bytes per token comes from. At 1M tokens that is about 0.9GB of global KV cache.

Mixed-precision weights: FP4 experts

The preview Flash card describes FP4 MoE experts with FP8 for most other parameters. The 0731 repo lists BF16, F8_E4M3, I8, and F32 tensors, so check which files you pull before sizing memory. Stacking another aggressive quant on top (an IQ3 GGUF, an fp8 KV cache) is where the garbage-output reports in the serving section came from.

Confirmed vs unconfirmed (Flash)

Confirmed first-party (model cards, release notes, arXiv:2606.19348, HF blog): 284B total / 13B active, 1M context, 384K recommended max output, CSA + HCA attention, 10% FLOPs / 7% KV cache vs V3.2, FP4/FP8 mixed weights in the preview, >32T training tokens, MIT. Not published for Flash: layer count, CSA/HCA layer split, and the lightning indexer's top-k.

DeepSeek V4 Flash Benchmarks: Self-Reported vs Independent

Most DeepSeek V4 Flash benchmark numbers come from DeepSeek. The preview model card reports SWE-bench Verified at 73.7 in non-thinking mode, 78.6 at high effort, and 79.0 at max. llm-stats repeats the 79.0 for DeepSeek-V4-Flash-Max (rank #17 of 116) and lists the 0423 preview at 78.6. For the 0731 GA checkpoint, DeepSeek reports DeepSWE 54.4, Terminal-Bench 2.1 82.7, and Toolathlon-Verified 70.3 on its own harness. The independent numbers are Vals.ai's SWE-bench Verified run, Code Arena votes, and the Artificial Analysis Intelligence Index.

DeepSeek V4 Flash benchmark scores by source (checked September 22, 2026)
BenchmarkFlash scoreSourceType
SWE-bench Verified73.7 / 78.6 / 79.0 (non-think / high / max)DeepSeek V4 Flash model cardSelf-reported
LiveCodeBench55.2 / 88.4 / 91.6DeepSeek V4 Flash model cardSelf-reported
GPQA Diamond71.2 / 87.4 / 88.1DeepSeek V4 Flash model cardSelf-reported
Terminal-Bench 2.1 (0731)82.7 (preview: 61.8)DeepSeek 0731 model cardSelf-reported
DeepSWE (0731)54.4 (preview: 7.3)DeepSeek 0731 model cardSelf-reported
Toolathlon-Verified (0731)70.3 (preview: 49.7)DeepSeek 0731 model cardSelf-reported
NL2Repo (0731)54.2 (preview: 39.4)DeepSeek 0731 model cardSelf-reported
SWE-bench Verified (0731, bash-only)88.8 (11th of 88)Vals.ai (archived September 1, 2026)Independent
Code Arena / WebDev (flash-high)1580 (tied 23rd)arena.ai, September 22, 2026Independent, human votes
AA Intelligence Index (0731, max)34 (median 18)Artificial AnalysisIndependent, standardized

The GA card also compares 0731 with outside models on DeepSeek's harness: Terminal-Bench 2.1 82.7 vs 81.0 for GLM-5.2 and 85.0 for Opus 4.8, and DeepSWE 54.4 vs 46.2 and 58.0. Those are useful for ranking, not for absolute numbers. No independent harness has re-run 0731's Terminal-Bench or DeepSWE claims as of September 22, 2026.

Vals.ai did run SWE-bench Verified on 0731 with a bash-only harness: 88.8, 11th of 88 models, against 96.4 for V4 Pro-0813 in 2nd. Vals then archived the benchmark on September 1, 2026 because scores had saturated, so V4.1 Flash will not get a Vals number. Scale's SWE-bench Pro leaderboard has not added an entry since July 9, 2026 and has no DeepSeek V4 result.

The independent signal is Artificial Analysis. Flash-0731 at max effort scores 34 on the Intelligence Index, against an 18 median for open-weight models of similar size. V4 Pro-0813 scores 36 and V4.1 Flash 39. AA also clocks Flash-0731 at 220.8 output tokens per second on DeepSeek's API, against a 67.6 median. For the V4 Pro side of this story, see the DeepSeek V4 guide.

Cheap Tokens, Expensive Habits

Price per token does not decide whether Flash is cheap. Output tokens per task do. Artificial Analysis measured Flash-0731 emitting 240M output tokens to run its full Intelligence Index, against a 140M median for open-weight models of similar size. V4 Pro-0813 used 160M. The smaller model reasons longer to reach its answers.

Output cost to run the AA Intelligence Index once (AA token counts x output price, September 22, 2026)
Model and rateAA Index scoreOutput tokensOutput price / 1MIndex output cost
V4 Flash 0731 on Morph (morph-dsv4flash)34240M$0.399625$95.91
V4.1 Flash, DeepSeek API off-peak39250M$0.60$150.00
V4.1 Flash on Morph (morph-dsv41flash)39250M$0.8$200.00
V4 Pro 0813, DeepSeek API off-peak36160M$1.98$316.80

This counts output tokens only and ignores input and cache, so read it as a ratio. Pro is terser, but it still costs more per Index run than either Flash at these rates. Within Flash, your effort setting is the main cost lever. 0731 has three levels (low / high / max) and AA ran max; pick the lowest level that clears your task. Then cache: on Morph, cached input on morph-dsv4flash is $0.0359375/M against $0.141953125/M uncached.

What makes a turn hit, how the cached rate compares across providers, and why routing a session across models resets the cache are covered on the prompt caching page.

Flash vs Pro: When Each Wins

Use Flash for high-volume and cost-sensitive work. Use Pro when a hard, long-horizon agent loop needs the extra active parameters. On Artificial Analysis the gap between Flash-0731 and Pro-0813 is 2 points (34 vs 36). Flash is also about 3x faster on DeepSeek's API: 220.8 vs 67.8 output tokens per second.

DeepSeek V4 Flash vs V4 Pro: the decision table (September 2026)
DimensionV4 Flash (0731)V4 Pro (0813)Who wins
Active params / token13B49BPro (capability)
AA Intelligence Index3436Pro (+2)
AA output speed (DeepSeek API)220.8 tok/s67.8 tok/sFlash (3.3x)
Output tokens to run the Index240M160MPro (terser)
SWE-bench Verified (max, self-reported)79.080.6Pro (+1.6)
SWE-bench Verified (Vals.ai, bash-only)88.896.4Pro (+7.6)
Self-host on a 128GB boxYes (quantized)No (multi-node)Flash
On DeepSeek's API todayRetired (routes to V4.1)Served, $1.98/M output off-peak, $3.96/M peakPro

The routing pattern most teams land on: Flash for bulk extraction, classification, first-draft edits, and background agent tasks; Pro for the hardest planning and multi-file refactors. Both use the same API shape and a 1M context, so switching is a model-string change.

Serving Footguns (vLLM / SGLang)

Flash gets its efficiency from low-precision attention and weights, and that is where self-hosted serving breaks. The serving flags are --tool-call-parser deepseek_v4 and --reasoning-parser deepseek_v4. Status of the filed issues as of September 22, 2026:

  • FP8 crashes under DP + EP (vLLM #43648), still open. DeepSeek-V4-Flash-FP8 on vLLM v0.20.1 with data-parallel + expert-parallel on H200 crashes after processing part of a benchmark run.
  • DSML tool parser mishandles wrapped/reserved args (vLLM #41240), closed May 6, 2026. Wrapper params named arguments or input were unwrapped when they were real schema fields. Upgrade past the fix if your tools use those names.
  • FP8 + pipeline-parallel garbage output (SGLang #25662), closed May 19, 2026. DeepSeek-V4-Flash-FP8 with --pp-size 8 on 8x H20 emitted garbage CJK tokens. Older SGLang builds still have it.
  • llama.cpp quantized KV cache (PR #25202), merged July 7, 2026. Builds before that date mishandle a quantized KV cache on DeepSeek V4.
  • V4.1 Flash is a separate code path. vLLM registers it as DeepseekV41ForCausalLM. One September 18 report (vLLM #57469) hit a host-memory OOM loading its ~510GB of weights on 4x H20 with 587GiB of node RAM. Budget host RAM, not just HBM, if you move up.
Practical serving takeaway

Validate your exact quant-plus-parallelism combination on a small eval before production. Keep the reasoning content in history: DeepSeek's encoding scripts retain reasoning_content that a naive Jinja template would strip, which breaks multi-turn tool calling.

Pricing and How to Run It on Morph

DeepSeek no longer sells V4 Flash 0731 by name. Its pricing page lists deepseek-flash (V4.1 Flash) at $0.15/M input (cache miss), $0.003/M (cache hit), and $0.60/M output off-peak, doubling to $0.30 / $0.006 / $1.20 during peak hours: 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays. The new rates took effect at 04:00 UTC on September 10. The legacy deepseek-v4-flash name bills at those rates. Artificial Analysis still lists Flash-0731 at $0.44/M input and $1.32/M output, the peak rate before September 10.

OpenRouter lists deepseek/deepseek-v4-flash-0731 from $0.04/M input and $0.64/M output (lowest-priced host, September 22, 2026).

$0.141953125 / $0.399625
morph-dsv4flash input / output per 1M tokens, $0.0359375 cached input, 1M context

Morph serves morph-dsv4flash on the 0731 checkpoint at 16-bit (bf16) activations with no fp8 activation quantization, at $0.141953125/M input and $0.399625/M output, flat at every hour, with a 1M context. For coding agents, Morph adds codegen-tuned speculative decoding and custom inference kernels. Call it through the OpenAI-compatible API:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.morphllm.com/v1",
    api_key="YOUR_MORPH_API_KEY",
)

resp = client.chat.completions.create(
    model="morph-dsv4flash",
    messages=[
        {"role": "user", "content": "Refactor this function to be async."},
    ],
)
print(resp.choices[0].message.content)

See Morph Open Source Models for the full lineup and pricing for every model's rates. For the broader DeepSeek V4 family and Claude Code setup, see the DeepSeek V4 guide and DeepSeek API guide.

DeepSeek V4 Flash API: Pricing and How to Call It

The DeepSeek V4 Flash API is OpenAI-compatible, so calling it is a base-URL and model-id change. The model id decides which checkpoint you get. On DeepSeek, deepseek-v4-flash now answers with V4.1 Flash. On Morph, morph-dsv4flash answers with the 0731 weights at $0.141953125/M input and $0.399625/M output with the full 1M context.

$0.141953125 / $0.399625
Morph input / output per 1M tokens
1M
Context window on Morph
Bearer auth
Authorization header, OpenAI-compatible

Call the DeepSeek V4 Flash API with the OpenAI SDK

Point any OpenAI-SDK client at https://api.morphllm.com/v1, pass your key as a Bearer token, and set the model to morph-dsv4flash. The raw HTTP call is a POST to https://api.morphllm.com/v1/chat/completions with an Authorization: Bearer YOUR_MORPH_API_KEY header. Flash is verbose, so set the reasoning effort to the lowest level that clears your task, and keep the model's reasoning_content in history so multi-turn tool calling does not break.

curl https://api.morphllm.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_MORPH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "morph-dsv4flash",
    "messages": [{"role": "user", "content": "Summarize this diff."}]
  }'

DeepSeek V4 Flash API pricing compared

DeepSeek V4 Flash API pricing by provider and checkpoint (September 22, 2026)
ProviderModel idCheckpointInput / 1MCache-hit inputOutput / 1M
Morphmorph-dsv4flashV4 Flash 0731$0.141953125 (flat)$0.0359375$0.399625 (flat)
Morphmorph-dsv41flashV4.1 Flash$0.12 (flat)$0.008$0.8 (flat)
OpenRouter (lowest host)deepseek/deepseek-v4-flash-0731V4 Flash 0731$0.04varies$0.64
DeepSeek (off-peak)deepseek-flash / deepseek-v4-flashV4.1 Flash$0.15$0.003$0.60
DeepSeek (peak)deepseek-flash / deepseek-v4-flashV4.1 Flash$0.30$0.006$1.20

Sources: DeepSeek's Models and Pricing page and change log; OpenRouter's public model list. Morph rates are the canonical pricing.

Getting a DeepSeek V4 API key on Morph

Sign up at morphllm.com, create a key in the dashboard, and send it as Authorization: Bearer YOUR_MORPH_API_KEY. One key covers the whole model lineup, including both Flash checkpoints. For the DeepSeek API walkthrough (setup, cache pricing, Claude Code), see the DeepSeek API guide.

Flash vs GLM 5.2 and GLM-5.3-Flash, MiniMax M3, Qwen 3.5 and Qwen3.8

Among open-weight coding models, V4 Flash 0731 is the one that fits a 1M context on a single 128GB box. Its successor V4.1 Flash scores higher but needs a full GPU node. DeepSeek's own 0731 card has GLM-5.2 at 81.0 on Terminal-Bench 2.1 against Flash-0731's 82.7, and 46.2 on DeepSWE against 54.4. Both GLM 5.2 and Qwen 3.5 now have successors. Z.ai shipped GLM-5.3 on August 18 and GLM-5.3-Flash on August 26, 2026. Alibaba's current line is Qwen3.8.

Cheap open-weight Flash-tier models vs DeepSeek V4 Flash (vendor docs and Artificial Analysis, September 22, 2026)
ModelSize (total / active)LicenseVendor API input / output per 1MAA Index
DeepSeek V4 Flash 0731284B / 13BMITretired on DeepSeek; routes to V4.134
DeepSeek V4.1 Flash~763B (552B backbone) / 8B-16BMIT$0.15 / $0.60 off-peak39
GLM-5.3-Flash320B / 18BMIT$0.15 / $0.5042
MiniMax-M3~428B / ~23Bminimax-community$0.30 / $1.20 (up to 512K input)29
Qwen3.8-Flash (hosted; Flash-Next weights)~180B / 6Bqwen-community-1.0$0.15 / $0.4740

GLM-5.3-Flash is the closest match: MIT weights, 1M context, image and video input, and the highest Index score of the group at the same $0.15 input price as V4.1 Flash. MiniMax M3 also reads images and video but trails on the Index. Qwen 3.5 ships under Apache 2.0, and so does the dense Qwen3.8-27B. The MoE Qwen3.8 weights use custom licenses.

Deep dives: GLM 5.2, GLM-5.3-Flash, MiniMax M3, and Qwen 3.5. For a ranked list, see the best open-source coding models in 2026.

Pros and Cons

Strengths
  • Low per-token price for the 0731 weights ($0.141953125/$0.399625 per M on Morph; from $0.04/$0.64 on OpenRouter)
  • Fits on a 128GB machine quantized (103GB IQ3 GGUF); V4.1 Flash does not
  • MIT license; preview and 0731 weights on Hugging Face
  • AA Intelligence Index 34 against an 18 median for its size, 2 points behind V4 Pro-0813
  • 220.8 output tokens/s on DeepSeek's API, against a 67.6 median
  • Pinned checkpoint: evals and prompts tuned on 0731 stay valid on hosts that keep serving it
Limitations
  • Retired on DeepSeek's API: deepseek-v4-flash routes to V4.1 Flash since September 10, 2026
  • Verbose: 240M output tokens on the AA Index vs a 140M median
  • 5 points behind V4.1 Flash on the AA Index (34 vs 39)
  • 0731 agentic gains (Terminal-Bench 82.7, DeepSWE 54.4) are self-reported
  • Text only; V4.1 Flash reads images
  • FP8 under vLLM DP + EP still crashes on H200 (vLLM #43648, open)

When to Use It, When Not

Use DeepSeek V4 Flash 0731 when you have evals, prompts, or fine-tunes built on it and need the same weights answering tomorrow. Use it when you self-host on one 128GB machine, or when output price matters more than 5 Index points: bulk extraction, classification, first-draft edits, and background agent tasks.

Move to V4.1 Flash when you want the higher scores or image input and can re-run your evals. Read the V4.1 Flash launch post first. Reach for V4 Pro on long-horizon agent loops where terseness matters. The self-hosting numbers for Pro are on the DeepSeek V4 Pro serving benchmarks section. On any task where Flash's verbosity blows your token budget, cap it at the low or high reasoning effort level and lean on caching.

Frequently Asked Questions

What is DeepSeek V4 Flash?

The smaller DeepSeek V4 model, previewed April 24, 2026 and GA on July 31, 2026 as DeepSeek-V4-Flash-0731, under MIT. 284B total / 13B active, 1M-token context, 384K recommended max output. DeepSeek retired it on its own API on September 10, 2026 in favor of V4.1 Flash. The weights stay on Hugging Face, and third-party hosts still serve them.

Does deepseek-v4-flash still call V4 Flash on DeepSeek's API?

No. Since September 10, 2026 the name is temporarily routed to V4.1 Flash and billed at the Flash price, and so is deepseek-v4-flash-vision-exp. The new name is deepseek-flash. To keep the 0731 weights, call deepseek/deepseek-v4-flash-0731 on OpenRouter or morph-dsv4flash on Morph.

What is the difference between DeepSeek V4 Flash and V4.1 Flash?

V4 Flash is 284B / 13B active, text only. V4.1 Flash is a causal encoder-decoder with a 552B backbone plus 196B of Engram memory and a vision encoder, about 763B total on Hugging Face. It activates 8B for input and 16B for output, reads images, and uses about a quarter of the KV cache. DeepSWE v1.1: 54.4 vs 74.2 (DeepSeek). AA Intelligence Index: 34 vs 39.

How big is DeepSeek V4 Flash?

284B total parameters, 13B active (Hugging Face's counter shows 304B for the 0731 repo). Unsloth's GGUFs are 103GB (UD-IQ3_XXS, at least 110GB RAM) to 162GB (UD-Q8_K_XL). V4.1 Flash has a 552B backbone and about 763B parameters in total, about 510GB on disk.

How is Flash different from V4 Pro?

Pro is 1.6T / 49B; Flash is 284B / 13B. AA scores Pro-0813 at 36 and Flash-0731 at 34, and measured Flash at 220.8 output tok/s vs Pro's 67.8. Pro used 160M output tokens to run the Index, Flash 240M. DeepSeek still serves Pro at $0.66/M input and $1.98/M output off-peak, or $1.32/M and $3.96/M at peak. On Vals.ai's bash-only SWE-bench Verified run, Pro-0813 scored 96.4 and Flash-0731 88.8.

What is DeepSeek V4 Flash's SWE-bench score?

73.7 (non-thinking), 78.6 (high), and 79.0 (max) on SWE-bench Verified, all from DeepSeek's model card. llm-stats lists Flash-Max at 79.0, rank #17. The independent number is Vals.ai: 88.8 for 0731 on a bash-only run, 11th of 88, before Vals archived the benchmark on September 1, 2026. For 0731, DeepSeek also reports DeepSWE 54.4 and Terminal-Bench 2.1 82.7 on its own harness. No independent tracker has re-run those two.

Does cheap Flash pricing mean cheap tasks?

Not automatically. AA measured Flash-0731 emitting 240M output tokens to run its Index against a 140M median for its size. Cap the reasoning effort level and cache to control it.

Can DeepSeek V4 Flash run locally?

Yes. Unsloth's UD-IQ3_XXS quant is 103GB and needs at least 110GB RAM. Hacker News users report running it under 128GB, including on a 128GB MacBook. Use a llama.cpp build after July 7, 2026 (PR #25202) if you quantize the KV cache.

Is there a free DeepSeek V4 Flash API?

The weights are free under MIT. The hosted APIs are paid. As of October 7, 2026, OpenRouter still resolves a deepseek/deepseek-v4-flash:free slug (the 0423 preview), but it lists zero live endpoints, so it serves no requests. OpenRouter's paid slugs are deepseek/deepseek-v4-flash (0423), deepseek/deepseek-v4-flash-0731, and deepseek/deepseek-v4.1-flash. DeepSeek bills deepseek-v4-flash at V4.1 Flash rates ($0.15/M input, $0.60/M output off-peak).

What serving flags does Flash need on vLLM/SGLang?

--tool-call-parser deepseek_v4 and --reasoning-parser deepseek_v4, with an FP8 KV cache via --kv-cache-dtype fp8. vLLM #41240 and SGLang #25662 are fixed; the FP8 DP + EP crash (vLLM #43648) is still open.

Is DeepSeek V4 Flash open source?

Yes. The preview and 0731 weights are on Hugging Face under MIT. DeepSeek ships Python encoding scripts instead of a Jinja chat template. The technical report is arXiv:2606.19348.

When were deepseek-chat and deepseek-reasoner retired?

July 24, 2026, 15:59 UTC. Since September 10, 2026, deepseek-v4-flash itself routes to V4.1 Flash. Current DeepSeek model names are deepseek-flash and deepseek-v4-pro.

Operator answer

The efficiency default for coding agents

Use it when cost per task and interactive speed matter more than winning the hardest frontier task.

Best fits

  • โœ“ Interactive coding agents
  • โœ“ High volume routine work
  • โœ“ Long sessions with strong prompt reuse

Escalate or test carefully

  • โ€ข Kernel engineering
  • โ€ข Real world robotics
  • โ€ข Temporal reasoning
  • โ€ข Spatial reasoning
Serving architecture

Cache the full agent session

Long coding sessions reuse system prompts, repository context, tool output, and prior turns. A useful production stack tiers that cache across GPU memory, CPU memory, and NVMe. GPU only cache sizing misses much of the cost per task opportunity.

Hardware guidance

Start with B300 when user speed matters, or B200 for aggregate throughput

Choose hardware around required speed per active user, then measure total capacity inside that latency target. Large batch throughput alone can hide a slow agent experience.

Related Articles

Dedicated inference planner

Plan a DeepSeek V4 Flash endpoint

Turn your team size and agent workload into a capacity estimate. Then validate the recommendation with your own traces.

Workload economics
Per model decision
900M
tokens per month
100 tok/s
required generation
$63
serverless per month
$14,016
dedicated per month
Sizing review required

An exact Morph capacity measurement is required before recommending a dedicated plan.

B200 is the compatible public platform. Dedicated capacity is invoiced monthly at the beginning of the month. Tokens are not billed separately.

Difference from serverless: $13,953 more per month.

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.

Talk to us about a private deployment

Run DeepSeek V4 Flash on Morph's OpenAI-Compatible API

morph-dsv4flash (V4 Flash 0731) at $0.141953125/M input and $0.399625/M output with a 1M context, or morph-dsv41flash (V4.1 Flash) at $0.12/$0.8. Pair either with WarpGrep so the context fills with the right code. WarpGrep runs $0.8 per 100K tokens.

Sources