---
title: "DeepSeek V4.1 Flash on the Morph API"
url: "https://www.morphllm.com/blog/deepseek-v4-1-flash"
description: "DeepSeek V4.1 Flash is live on Morph as morph-dsv41flash: a 552B-backbone MoE (about 763B total with Engram memory and vision) with 8B active for input and 16B for output, image input, native tool calling, thinking on by default with an adjustable budget, and a 1M context. $0.15 per million input, $0.01 cached, $0.60 output."
date: "2026-09-10"
author: "Tejas Bhakta"
---
# DeepSeek V4.1 Flash on the Morph API

DeepSeek V4.1 Flash is serving on Morph as of today. The model id is `morph-dsv41flash`, it sits behind the same OpenAI-compatible base URL as everything else we run, and it costs $0.15 per million input tokens, $0.01 per million cached input tokens, and $0.60 per million output tokens.

## What it is

V4.1 Flash is a mixture-of-experts with a 552B backbone, plus 196B of Engram conditional memory and a vision encoder, for about 763B parameters on Hugging Face. It activates 8B parameters per token on input and 16B on output. That ratio is the reason it lands in the "Flash" tier of the DeepSeek line. Decode only touches the active experts, so the per-token cost is that of a mid-sized model while the routing pool behind it is much larger.

It has a 1M-token native context, so a whole repository and a long agent trajectory fit without eviction.

It takes images. V4 Flash on Morph was text only; V4.1 Flash reads screenshots, diagrams, and rendered pages in the same `messages` array, using the standard `image_url` content part. Video input is not supported.

It calls tools natively. Pass `tools` the way you already do and you get structured `tool_calls` back. JSON mode and JSON schema structured outputs work as well.

And it thinks by default. Every request runs a reasoning pass before the answer unless you turn it off, and the length of that pass is a knob you control.

## The thinking budget

V4.1 Flash reads `reasoning_effort` on the request. The named tiers `low`, `high`, `xhigh`, and `max` are native to the model, and an integer from 1 to 100 sets the thinking budget directly. `high`, which is what the model uses when you send nothing, corresponds to 50. `none` turns thinking off for that request.

Two tiers from the OpenAI vocabulary are not in the model's own: we map `medium` to `high` and `minimal` to `low`, so a client written against another provider keeps working without a code change.

In practice: leave it alone for agentic coding, set a small integer for classification and extraction where the thinking pass is wasted, and set `none` for anything you want streamed at full speed with no preamble.

## Pricing

<table>
  <thead>
    <tr><th></th><th>Per 1M tokens</th></tr>
  </thead>
  <tbody>
    <tr><td>Input</td><td>$0.15</td></tr>
    <tr><td>Cached input</td><td>$0.01</td></tr>
    <tr><td>Output</td><td>$0.60</td></tr>
  </tbody>
</table>

Cached input is 7% of the uncached rate. Prefix caching is on for every request with nothing to configure and no cache-write surcharge. When a request shares a prefix with earlier traffic (the system prompt, the tool definitions, the conversation so far), those tokens skip prefill and bill at $0.01. For a coding agent, where the prompt is mostly the same file and the same tools turn after turn, that is most of the bill.

No per-seat fees, no minimums. Every response reports how much of the prompt was served from cache in `usage.prompt_tokens_details.cached_tokens`.

## Using it

It is one endpoint and one model string.

```bash
curl https://api.morphllm.com/v1/chat/completions \
  -H "Authorization: Bearer $MORPH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "morph-dsv41flash",
    "messages": [
      {"role": "user", "content": "Find the bug in this function and explain it in two sentences.\n\ndef median(xs):\n    xs.sort()\n    return xs[len(xs) // 2]"}
    ]
  }'
```

The Python client is the stock OpenAI SDK. This one sends an image and caps the thinking budget:

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.morphllm.com/v1",
    api_key=os.environ["MORPH_API_KEY"],
)

res = client.chat.completions.create(
    model="morph-dsv41flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is wrong with the layout in this screenshot?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/render.png"}},
        ],
    }],
    extra_body={"reasoning_effort": 25},
)

print(res.choices[0].message.content)
print(res.usage.prompt_tokens_details.cached_tokens)
```

Swap `extra_body={"reasoning_effort": "none"}` in and the same call skips the thinking pass. Add `tools=[...]` and it calls them.

There is also an Anthropic-compatible `/v1/messages` endpoint, so Claude Code can run on it by setting `ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN`, and `ANTHROPIC_MODEL`. The [Claude Code guide](https://docs.morphllm.com/guides/claude-code) has the exact lines.

## Why Morph

The price on this model comes from two decisions about how we serve it, not from a discount.

The first is the hardware. V4.1 Flash runs on 8xB200 spot nodes. Spot capacity costs a fraction of on-demand, and the fleet autoscales across zones, so when a node gets reclaimed the model follows the capacity instead of going down with the node. You get B200 economics without paying B200 reservation prices.

The second is the cache. Each worker keeps a three-tier prefix cache: the hot set in GPU memory, a much larger tier in host RAM, and a local SSD tier behind that sized to the node's disk. A request whose prefix lives in any of the three skips the prefill for those tokens. Cached input at $0.01 is not a marketing rate; it is what those tokens cost us to serve once they are already resident.

Put the two together and the model an agent hits a thousand times a day, with the same system prompt and the same tools on every call, is cheap in the place agents actually spend tokens.

The [full lineup](/products/models) sits behind the same base URL, and live per-token rates for every model are at [/api/models/json](/api/models/json). Grab a key from the [dashboard](/dashboard) and change one string.
