---
title: "Morph x AWS: The Fastest Code-Editing Model"
url: "https://www.morphllm.com/blog/morph-aws-case-study"
description: "AWS selected Morph as a featured case study — covering the infrastructure behind 10,000 tokens/sec, how Binance achieved 50-70% productivity gains, and why we built on EC2 P5 and SageMaker."
date: "2026-03-05"
author: "Tejas Bhakta"
---
# Morph x AWS: The Fastest Code-Editing Model

![AWS and Morph](/images/aws-case-study.png)

**[Read the AWS case study](https://aws.amazon.com/solutions/case-studies/morph-case-study/)**

Anyone who's built with a coding agent has watched this happen. Claude reads the repo, works out the refactor, writes it. The thinking part is quick. Then apply takes over and the whole thing bogs down.

So you wait. Sometimes two minutes, while another process shoves that output back into your files. Often it fails. Imports drop, the indentation goes sideways, or the diff lands on the wrong function and you don't notice until the tests do.

Morph exists to delete that second half. AWS put out a writeup on it last week.

## The case study

These aren't ghostwritten. AWS finds a company that fixed something real, then documents the problem and the infrastructure calls and the numbers behind it. Ours went live last week.

The writeup traces our path from 1,000 tokens/sec up to 10,000, names the AWS services that carried us there, and gets into what changes for a team once AI coding agents are in front of real users.

## The numbers

| What | Before | Now |
|------|--------|-----|
| Throughput | 1,000 tok/sec | **10,000 tok/sec** |
| 15k-token multifile refactor | 20 seconds | **under 400ms** |
| Single-file edit | 2-5 minutes | **under 1 second** |

Once apply runs at 10,000 tok/sec you stop thinking about it. It's not a step you wait on anymore. It just happens, and your attention never leaves the code. That was the target the whole time.

Binance shipped it and measured. Their teams came back with 50-70% productivity gains. Real number, real customer, named in the study. We didn't anonymize it.

## The infrastructure

Training ran on EC2 P5 boxes. H100s. Memory bandwidth is the thing that makes speculative decoding worth doing at this scale, and the H100 has it. We wrote our own CUDA kernels that fuse attention and feedforward into a single pass, which drops a bunch of memory roundtrips and gets us to 2.1TB/s per GPU.

Deployment runs on SageMaker. Not the default choice, a deliberate one. Enterprise customers need to know their code never leaves their boundary, and SageMaker's isolation handed us that guarantee. We didn't have to build a security layer ourselves.

Here's a line from the study:

> "AWS is infrastructure I can trust... AWS has tried-and-tested solutions, and I'm not going to encounter hardware failures." — Tejas Bhakta, Founder and CEO

Hardware reliability sounds boring until a run dies at hour 18 of 20. Then it's not boring. You've burned the compute and every customer waiting on the next checkpoint slips a day.

## On AWS Marketplace

You can pull Morph in through AWS Marketplace now. Procurement you already have, billing you already have. No new vendor to stand up.

Already on AWS? Then there's nothing to set up.

[Read the full case study](https://aws.amazon.com/solutions/case-studies/morph-case-study/) or [grab an API key](/dashboard/api-keys) and benchmark it yourself.

---

**Building something with coding agents?** [Reach out](mailto:info@morphllm.com). We work directly with teams running Morph at scale.
