---
title: "Fast Apply: Merge LLM code edits at 10,500 tok/s"
url: "https://www.morphllm.com/products/fastapply"
canonical_url: "https://www.morphllm.com/products/fastapply"
docs_url: "https://docs.morphllm.com/sdk/components/fast-apply"
description: "Specialized model that applies LLM-generated code edits to existing files. 10,500 tok/s, 98% accuracy, OpenAI-compatible API. Use when a coding agent outputs a lazy snippet, patch, or partial edit that needs to be merged into the original file."
---

# Fast Apply

Applies LLM-generated code edits to existing files. Use when a coding agent outputs a lazy snippet, `// ... existing code ...` marker, patch, or partial edit and you need it merged into the original file without breaking semantics. Models: `morph-v3-fast` (7B, 10,500 tok/s) and `morph-v3-large` (14B, 2,600 tok/s). OpenAI-compatible.

- Canonical: https://www.morphllm.com/products/fastapply
- Docs: https://docs.morphllm.com/sdk/components/fast-apply
- Model spec: https://docs.morphllm.com/models/apply
- API reference: https://docs.morphllm.com/api-reference/endpoint/apply

## Problem

Coding agents produce edits as fragments, not full files. Full-file rewrites waste tokens and introduce unrelated changes. Unified diffs fail on fuzzy context. Search-and-replace breaks on whitespace or quote differences. Fast Apply merges a minimal "lazy" edit into an original file deterministically.

## Why it's fast

A small (7B/14B) model trained end-to-end for a single task (apply), served on a custom inference engine with hand-written GPU kernels tuned for the apply workload. Not a fine-tune on vLLM/TGI. The serving stack was built for this workload specifically, which is how the same parameter count hits 10,500 tok/s.

## Quickstart

```ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.morphllm.com/v1",
  apiKey: process.env.MORPH_API_KEY,
});

const response = await client.chat.completions.create({
  model: "morph-v3-fast",
  messages: [
    {
      role: "user",
      content: `<code>${originalFile}</code>\n<update>${lazyEdit}</update>`,
    },
  ],
});

const merged = response.choices[0].message.content;
```

Input format: the original file inside `<code>` and the lazy edit inside `<update>`. Fast Apply returns the fully merged file. See the [prompting guide](https://docs.morphllm.com/guides/prompting) and [XML vs JSON tool calls](https://docs.morphllm.com/guides/xml-tool-calls) for why XML beats JSON for editing.

## Models & pricing

| Model | Size | Speed (tok/s) | Input $/1M | Output $/1M | Use when |
|-------|------|---------------|------------|-------------|----------|
| `morph-v3-fast` | 7B | 10,500 | 0.80 | 1.20 | Real-time IDE, streaming, agent inner loop |
| `morph-v3-large` | 14B | 2,600 | 0.90 | 1.90 | Higher-accuracy batch edits, complex multi-hunk |

Context window: 32K tokens. Sub-10ms overhead, sub-second cold starts. Full pricing: https://www.morphllm.com/pricing

## When to use

- Agent generated a lazy snippet and you need the full file back
- Streaming edits from GPT-4/Claude/Gemini into a real file
- Multi-hunk edits to a single file in one call
- Replacing brittle `str_replace_editor` / unified-diff post-processing

## When not to use

- Greenfield file generation (use the planning model directly)
- Single-character edits (cheaper to do client-side)
- Binary files, minified JS, or non-textual content

## Integrations

- [Claude Code](https://docs.morphllm.com/guides/claude-code): speed up Claude Code's file edits
- [Vercel AI SDK](https://docs.morphllm.com/guides/ai-sdk): stream via `morph:morph-v3-fast`
- [Agent Tools / edit_file](https://docs.morphllm.com/guides/agent-tools): wire Fast Apply behind an `edit_file` tool call
- [One-shot edit_file prompt](https://docs.morphllm.com/guides/oneshot): drop-in edit_file implementation
- [MCP server](https://docs.morphllm.com/mcpquickstart): Cursor, Claude Desktop, Windsurf, Cline

## Deployment

- Real-time (hosted): `api.morphllm.com`, sub-10ms overhead, global
- [Self-hosted](https://docs.morphllm.com/api-reference/self-hosting): on-prem / air-gapped
- [Enterprise Apply](https://docs.morphllm.com/api-reference/endpoint/enterprise): custom model configurations
- [Report API](https://docs.morphllm.com/api-reference/endpoint/report): report failed merges for model improvement

## See also

- [Quickstart (first apply in under 2 minutes)](https://docs.morphllm.com/quickstart)
- [TypeScript SDK](https://docs.morphllm.com/sdk/quickstart)
- [API reference: apply endpoint](https://docs.morphllm.com/api-reference/endpoint/apply)
- [Context selection guide](https://docs.morphllm.com/guides/context)
- [Authentication](https://docs.morphllm.com/auth)
- Blog: [Fast Apply and Fast Agents](https://www.morphllm.com/blog/fast-apply-fast-agents)
- Blog: [Diffs vs Fast Apply](https://www.morphllm.com/blog/diffs-vs-fast-apply)
- Blog: [Morph breaks the 10k tok/s barrier](https://www.morphllm.com/blog/morph-breaks-10k-barrier)
- Benchmarks: https://www.morphllm.com/benchmarks/fast-apply
