Grok Build is xAI's terminal coding agent, open source under Apache-2.0 since July 14, 2026, at CLI version 1.0.24 and 26,742 GitHub stars. It runs grok-4.6 on a 500K context at $2 per million input tokens and $6 per million output tokens. Its subagent coordinator ships a default concurrency cap of 32, not the 8 that launch coverage reported. Claude Code is at v2.1.271, published September 14, 2026, and runs Claude Opus 5, Sonnet 5, Fable 5.1 and Haiku 4.5, the first three on 1M token context windows, bundled into Claude Pro at $20 per month and Max from $100. On CursorBench 3.2, the one benchmark both vendors publish, Anthropic reports 73.4% for Fable 5.1 and 70.0% for Opus 5; xAI reports 69.9% for Grok 4.6. Choose Claude Code for context depth and a 33-event hook system. Choose Grok Build for an Apache-2.0 harness you can fork and cheaper tokens.
The two agents converged. Grok Build reads CLAUDE.md and .claude/settings.json hooks directly. Both ship plan modes, MCP, skills, worktrees, headless -p, and subagents. What separates them now is the model, the license, the concurrency ceiling, and the price per token.
Quick Verdict
- Choose Grok Build if: You want an Apache-2.0 harness you can read and fork, cheaper tokens at $2 in and $6 out per million, and up to 32 concurrent subagents by default
- Choose Claude Code if: You need a 1M token context, 33 hook lifecycle events, and the highest CursorBench 3.2 score published by either vendor at 73.4%
- Run both if: Your repo already has an
AGENTS.mdorCLAUDE.md. Grok Build loads Claude Code's rules and hook files as-is, so the switching cost is one install command
| Feature | Grok Build | Claude Code |
|---|---|---|
| Developer | xAI | Anthropic |
| Version | CLI 1.0.24 | v2.1.271 |
| Released | Beta May 14, 2026; open source Jul 14, 2026 | May 2025 |
| License | Apache-2.0 (harness) | Proprietary |
| Default model | grok-4.6 | Claude Opus 5 |
| Context window | 500K (grok-4.6) | 1M (Opus 5, Sonnet 5, Fable 5.1) |
| API price per M tokens | $2 in / $6 out | $5 in / $25 out (Opus 5) |
| Subagent concurrency | 32 by default, configurable | Task tool, no published cap |
| CursorBench 3.2 | 69.9% (Grok 4.6) | 73.4% Fable 5.1, 70.0% Opus 5 |
| Project memory | AGENTS.md, also reads CLAUDE.md | CLAUDE.md |
| MCP Support | Yes | Yes |
| Hooks | 14 events | 33 events |
| Headless Mode | Yes (-p flag) | Yes (-p flag) |
| Agent Protocol | ACP (full support) | Subagent spawning |
| Maturity | Open source, 26,742 stars | 2+ years production use |
Grok Build vs Claude Code by Version, Release Date, and Price (September 2026)
Both agents ship several times a week, which is why comparisons written in spring 2026 describe products that no longer exist. The table records what each tool was at when this page was last checked, so you can measure how far it has moved since.
| Tool | Version | Released | Open source | Subscription price |
|---|---|---|---|---|
| Grok Build | CLI 1.0.24 | Sep 9, 2026 (monorepo sync) | Yes, Apache-2.0 | SuperGrok or X Premium+ (price not verifiable) |
| Claude Code | v2.1.271 | Sep 14, 2026 | No | Claude Pro $20/mo, Max from $100/mo |
| OpenAI Codex CLI | rust-v0.154.0 | Sep 9, 2026 | Yes (CLI) | ChatGPT Plus $20/mo, Pro from $100/mo |
| Model | Context | Input / M tokens | Output / M tokens | Released |
|---|---|---|---|---|
| grok-4.6 (under 200K prompt) | 500K | $2.00 | $6.00 | Aug 12, 2026 |
| grok-4.6 (200K and above) | 500K | $4.00 | $12.00 | Aug 12, 2026 |
| grok-build-0.1 | 256K | $1.00 | $2.00 | Not published |
| Claude Opus 5 | 1M | $5.00 | $25.00 | Jul 24, 2026 |
| Claude Fable 5.1 | 1M | $10.00 | $50.00 | Sep 1, 2026 |
| Claude Sonnet 5 | 1M | $2.00 | $10.00 | Not published |
| Claude Haiku 4.5 | 200K | $1.00 | $5.00 | Oct 2025 |
Two things fell out of the version check. Grok Build has no git tags and no GitHub releases; the only version marker is the xai-grok-version crate, whose description reads "Lockstepped grok CLI version" and which sat at 1.0.24 on the September 9, 2026 sync commit. And xAI ships a second, cheaper coding model, grok-build-0.1, at $1 in and $2 out per million tokens on a 256K context. It is priced in the public model table but is not the Grok Build default.
Architecture: Parallel Breadth vs Reasoning Depth
The fundamental difference between these two tools is architectural, and it shapes everything else.
Grok Build: Plan, Then Fan Out
Grok Build cycles three session modes with Shift+Tab: plan, auto, and always-approve. Plan mode gates file edits so only the session plan file can be written until you approve it, and that gate is independent of the permission mode. You enter it with /plan and reopen a plan with /view-plan.
Once the plan is approved, the main agent delegates to subagents. There are three built-in types: general-purpose with full capability, explore which can read, list and search but cannot run a shell or edit, and plan which drafts an implementation plan with the same restrictions. You add or override types under .grok/agents/.
The tradeoff is context. Each subagent is an independent child session with its own context that returns a summary to the parent when it finishes. For a task where the relationship between distant files matters, no single agent sees the whole picture.
Claude Code: One Agent, Deep Context
Claude Code takes the opposite default. One agent on a 1M token context window that can hold most repositories at once. Instead of fanning out first, it reasons within a single context, tracking dependencies across files and producing changes that are internally consistent.
Claude Code can spawn subagents, but this is opt-in rather than default. The primary workflow is sequential: understand the full picture, plan the change, execute across files. This produces more deterministic output and takes longer per task.
Grok Build: Parallel Breadth
Up to 32 concurrent subagents by default, in three built-in types. Each is an independent child session that returns a summary. Higher throughput on independent work, fragmented context on coupled work.
Claude Code: Reasoning Depth
One agent on a 1M token context. Cross-file dependency tracking, deterministic multi-file edits, and 33 hook events to gate the loop. Subagents are available but off the default path.
Parallel breadth wins on tasks that decompose cleanly into independent units: migrating 40 files to a new API, writing tests per module, sweeping a lint rule. Reasoning depth wins on tasks with one correct answer that depends on interconnections: refactors across coupled modules, bugs in deeply nested call chains, schema migrations that touch billing and auth at once.
Multi-Agent Approach Comparison
Both tools support multi-agent workflows. The default behavior, the concurrency ceiling, and how you raise it differ.
Grok Build: 32 Concurrent Subagents, Not 8
Launch coverage in May 2026 put Grok Build at 8 parallel agents. The shipped source says otherwise. DEFAULT_MAX_CONCURRENT is 32 in the subagent admission module at crates/codegen/xai-grok-tools/src/implementations/grok_build/task/admission.rs, whose file comment reads "Session-scoped subagent spawn limits, enforced by the coordinator."
Two environment variables control it, and neither appears in the published settings reference at docs.x.ai. GROK_MAX_CONCURRENT_SUBAGENTS sets the cap, and the source notes a limit "can be adjusted but never disabled": a value of 0 is clamped to 1. GROK_SUBAGENT_LIMIT_BEHAVIOR chooses what happens to a spawn that arrives at the ceiling, accepting queue or fail. The default is queue, so a 33rd subagent waits rather than erroring, and any other value logs a warning and falls back to queue.
Because the default limit behavior is queue, a prompt that spawns 100 subagents does not fail loudly. It runs 32 at a time and holds the rest, which reads as a stall rather than a limit. Set GROK_SUBAGENT_LIMIT_BEHAVIOR=fail when you want the agent to be told it hit the ceiling. Neither variable is documented, so both were read from the Apache-2.0 source on September 14, 2026 and could change without a changelog entry.
Claude Code: Subagents On Demand
Claude Code's subagents are spawned explicitly when parallelism is needed. The primary agent coordinates, delegates specific investigation or implementation tasks, and synthesizes results. Each subagent gets its own context window and tool access, and SubagentStart and SubagentStop hooks fire around each one.
The key difference: Claude Code's subagents are coordinated by a primary agent that holds the full context in a 1M token window. Grok Build's subagents are more autonomous, each returning a summary to a parent working within a 500K window.
| Aspect | Grok Build | Claude Code |
|---|---|---|
| Default mode | Single agent, delegates on larger tasks | Single agent |
| Concurrency cap | 32 (GROK_MAX_CONCURRENT_SUBAGENTS) | No published cap |
| At the cap | Queues by default, or fails if configured | Not published |
| Built-in types | general-purpose, explore, plan | Custom agents in .claude/agents/ |
| Context sharing | Independent child session per subagent | Subagents inherit coordinator context |
| Lifecycle hooks | SubagentStart, SubagentStop | SubagentStart, SubagentStop, TaskCreated, TaskCompleted |
| Cancel behavior | cancel_subagents_on_turn_cancel: ask by default | Not published |
| Best for | Work that decomposes into independent units | Tasks requiring unified context |
Arena Mode, the automatic best-of-N evaluator that May 2026 coverage described, is gone from the documentation. The string "arena" does not appear anywhere in docs.x.ai, including its llms.txt index, and no /arena command exists in the Grok Build command table as of September 14, 2026. xAI has not published a note explaining the removal. The nearest current equivalent is manual: spawn several subagents on the same task and compare their summaries yourself.
Pricing Comparison
Token pricing is verifiable from both vendors. Subscription pricing is only verifiable on the Anthropic side: xAI's consumer pricing pages at x.ai and grok.com returned 403 and 404 respectively when checked on September 14, 2026, so the $299 per month SuperGrok Heavy figure this page carried from the May 2026 beta is marked unverified below rather than repeated as fact.
| Model | Input / M | Cached input / M | Output / M |
|---|---|---|---|
| grok-4.6 (under 200K prompt) | $2.00 | $0.50 | $6.00 |
| grok-build-0.1 | $1.00 | $0.20 | $2.00 |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 |
| Claude Sonnet 5 | $2.00 | Not listed here | $10.00 |
| Claude Fable 5.1 | $10.00 | Not listed here | $50.00 |
| Tier | Grok Build | Claude Code |
|---|---|---|
| Free | None | None (Claude Free excludes Claude Code) |
| Entry | SuperGrok (price unverified) | $20/mo Pro, or $17/mo billed annually |
| Heavy use | SuperGrok Heavy (price unverified) | Max from $100/mo, 5x or 20x Pro usage |
| Team | Not published | $25/mo per seat, $20 billed annually |
| Usage model | Shared weekly pool across Grok surfaces | Shared across Claude and Claude Code |
Claude's pricing page lists Max as "From $100 per month" and says you "choose 5x or 20x more usage than Pro." It does not display a separate figure for the 20x tier anywhere on the page, including the full feature comparison table, which lists Max 5x and Max 20x as columns with no prices. Comparisons quoting $200 per month for Max 20x are quoting a number that is not on the page as of September 14, 2026.
Cost Per Successful Task
Raw subscription price is only half the equation, and the multi-agent side of it cuts against Grok Build. Thirty-two concurrent subagents each carry their own context, so the token cost of a fanned-out task scales with the number of children, not with the size of the change. Claude Code's single-agent default spends more per turn on a bigger context and fewer turns overall.
The weekly pool is the practical constraint on the Grok side. A developer who built an open-source macOS-binary compatibility layer with Grok Build reported on Hacker News in August 2026 that despite using Grok 4.5 at medium effort with "no sub-agents, just standard chats and Plan Mode," they "hit the weekly usage limits all the time and had to sit around waiting for the cooldowns." That was a single-agent workload. Fanning out to 32 subagents draws on the same pool.
Grok Build vs Claude Code vs Codex
The three-way question comes up because all three are terminal agents from frontier labs, and two of the three are open source. Prices below are from each vendor's own pricing page on September 14, 2026.
| Grok Build | Claude Code | OpenAI Codex CLI | |
|---|---|---|---|
| Version | CLI 1.0.24 | v2.1.271 | rust-v0.154.0 |
| Last shipped | Sep 9, 2026 | Sep 14, 2026 | Sep 9, 2026 |
| License | Apache-2.0 | Proprietary | Open source (CLI) |
| Default model | grok-4.6 | Claude Opus 5 | GPT-5.6 family |
| Context window | 500K | 1M | Not published |
| Entry subscription | SuperGrok (unverified) | $20/mo Pro | $20/mo ChatGPT Plus |
| Cheapest paid tier | Not published | $20/mo Pro | $8/mo Go |
| Heavy tier | SuperGrok Heavy (unverified) | Max from $100/mo | Pro from $100/mo |
| CursorBench 3.2 | 69.9% (Grok 4.6) | 73.4% Fable 5.1, 70.0% Opus 5 | 67.2% (GPT-5.6 Sol) |
| Terminal-Bench 4.0 | Not published | 55.8% Fable 5.1, 52.3% Opus 5 | 37.3% (GPT-5.6 Sol) |
| Project memory | AGENTS.md | CLAUDE.md | AGENTS.md |
Codex is the cheapest way in. OpenAI sells a Go tier at $8 per month for lightweight coding, below the $20 Plus tier that includes web, CLI, IDE and iOS access, and its Pro tier starts at $100 per month for 5x or 20x higher rate limits than Plus. Read the full head-to-head on Codex vs Claude Code.
The Terminal-Bench 4.0 and CursorBench 3.2 numbers in that table for Codex are Anthropic's measurements of GPT-5.6 Sol, published alongside its own scores, not OpenAI's. xAI has not published a Terminal-Bench 4.0 number at all; its Grok 4.6 announcement reports Terminal-Bench v3.0 at 26%, up from 15.7% for Grok 4.5, which is a different harness version and does not compare.
Benchmarks
Benchmark data checked September 14, 2026. The comparison is harder than it was in May, because the two labs now publish on almost disjoint benchmark sets.
| Benchmark | Grok 4.6 (xAI) | Claude (Anthropic) |
|---|---|---|
| CursorBench 3.2 | 69.9% | 73.4% Fable 5.1 (max effort), 70.0% Opus 5 |
| Terminal-Bench | 26% (v3.0) | 55.8% Fable 5.1, 52.3% Opus 5 (v4.0) |
| SWE-bench Verified | Not published | Not published for Opus 5 |
| Benchmark | Score | Source |
|---|---|---|
| DeepSWE v1.1 | 65.9% (Grok 4.6, up from 54%) | xAI |
| FrontierCode v1.1 | 61.3% (Grok 4.6, up from 56.6%) | xAI |
| APEX-Agents | 57.5% (Grok 4.6, up from 47.1%) | xAI |
| Terminal-Bench-Science 0.1 | 52.6% Fable 5.1, 29.0% Opus 5 | Anthropic |
| RedlineBench | 57.0 Fable 5.1, up from 47.9 Fable 5 | Anthropic |
| Frontier-Bench v0.1 | Opus 5 more than doubles Opus 4.8, no absolute number | Anthropic |
Earlier versions of this page cited 80.8% on SWE-bench Verified for Claude Code. That number belongs to Claude Opus 4.6, measured at 80.84% over 25 trials and published in February 2026. Anthropic has not published a SWE-bench Verified score for Claude Opus 5 or Claude Fable 5.1; the Opus 5 announcement leads with Frontier-Bench v0.1, ARC-AGI 3, OSWorld 2.0 and CursorBench 3.2 instead. Treat any 2026 comparison still quoting 80.8% as a current Claude Code score as out of date.
For context, other terminal agents: OpenAI Codex runs GPT-5.6 Sol, which Anthropic measured at 67.2% on CursorBench 3.2 and 37.3% on Terminal-Bench 4.0. Gemini CLI now ships as Antigravity CLI on Google's shared Antigravity harness.
Context Window and Codebase Handling
Context management is where the architectural difference becomes most visible in daily use.
Claude Code: 1M Tokens, Single Agent
Claude Opus 5, Claude Sonnet 5 and Claude Fable 5.1 all carry 1M token context windows with 128K max output. On the current tokenizer, Anthropic puts 1M tokens at roughly 555,000 words. That is large enough to hold most repositories in a single context, so the agent reads the full structure, understands architectural patterns, and tracks dependencies across files. Long sessions still benefit from proactive compaction, and PreCompact and PostCompact hooks fire around it.
Grok Build: 500K Per Agent, Distributed
Grok Build's default model, grok-4.6, carries a 500K context window. Crossing 200K prompt tokens doubles the price to $4 in and $12 out per million. Each subagent is an independent child session with its own context that returns only a summary to the parent, so aggregate capacity across 32 children is large but no single agent sees the whole picture.
For large codebases with deeply interconnected modules, this is a meaningful tradeoff. Refactoring an auth module that touches billing, API routes, and database schemas works better when a single agent holds all four concerns at once. For feature additions confined to a directory, the distributed approach has less downside.
| Aspect | Grok Build | Claude Code |
|---|---|---|
| Max context per agent | 500K (grok-4.6) | 1M (Opus 5, Sonnet 5, Fable 5.1) |
| Long-context surcharge | 2x price above 200K prompt tokens | None published at 1M |
| Max output | Not published | 128K tokens, 300K on Batch API beta |
| Cross-file reasoning | Per subagent, summary to parent | Single unified context |
| Project memory | AGENTS.md, CLAUDE.md, .grok/rules/, .claude/rules/ | CLAUDE.md |
| Context management | /compact-mode, compaction transcript | PreCompact and PostCompact hooks |
Using Grok in Claude Code
The mechanism exists, and xAI has marked it for removal. Both halves matter, so here is the exact state as of September 14, 2026.
What xAI exposes
xAI serves a POST /v1/messages endpoint at https://api.x.ai/v1, documented as "compatible with the Anthropic API." Claude Code reads ANTHROPIC_BASE_URL to override the API endpoint and ANTHROPIC_AUTH_TOKEN as the value of the Authorization header, prefixed with Bearer. Wiring those two together is the whole configuration:
export ANTHROPIC_BASE_URL=https://api.x.ai/v1
export ANTHROPIC_AUTH_TOKEN=<your xAI API key>
export ANTHROPIC_MODEL=grok-4.6
claudeWhy it is not a setup to build on
xAI's own REST reference files that endpoint under "Legacy & Deprecated" and prints, three times on the page, "Deprecated: The Anthropic SDK compatibility is fully deprecated. Please migrate to the Responses API or gRPC." The /v1/complete endpoint carries the same notice. xAI has not published a shutdown date, so this is a documented deprecation without a documented end.
Claude Code also degrades on a non-first-party host, and its environment variable reference says so explicitly. Pointing ANTHROPIC_BASE_URL at a non-first-party host disables MCP tool search by default, recoverable with ENABLE_TOOL_SEARCH=true only if the endpoint forwards tool_reference blocks. As of v2.1.196, Remote Control is disabled whenever ANTHROPIC_BASE_URL points anywhere other than api.anthropic.com. Anthropic documents third-party routing for Amazon Bedrock, Google Cloud, Microsoft Foundry and LLM gateways, all serving Anthropic models. It does not document running another lab's model in Claude Code.
Grok Build reads Claude Code's configuration natively. Its project-rules loader reads AGENTS.md, AGENT.md, CLAUDE.md, Claude.md and CLAUDE.local.md, plus every Markdown file in .grok/rules/, with .claude/rules/ and .cursor/rules/ read for compatibility. Its hook loader reads .claude/settings.json and .cursor/hooks.json, and maps Claude tool names such as Bash, Read and Edit onto Grok's automatically. A repo configured for Claude Code runs in Grok Build with no migration work. The reverse, running Grok inside Claude Code, is the path xAI deprecated.
Ecosystem: MCP, Hooks, Plugins
Both tools support the same categories of extensibility. The hook surface is where the gap is measurable rather than impressionistic.
| Feature | Grok Build | Claude Code |
|---|---|---|
| MCP servers | Supported, with on-demand tool search | Supported, with MCP tool search |
| Hook events | 14 | 33 |
| Hook transport | Shell command or HTTP POST | Command, HTTP, prompt, agent, async |
| Blocking hook | PreToolUse only | PreToolUse and PermissionRequest |
| Plugins | Skills, plugins, marketplaces | Skills, plugins, custom commands |
| Project memory | AGENTS.md, also reads CLAUDE.md | CLAUDE.md |
| Agent protocol | ACP (full support) | Subagent spawning |
| Headless mode | Yes (-p flag) | Yes (-p flag) |
| Sandbox | off, workspace, read-only, strict | Sandboxed bash with per-command allowed_domains |
| Source available | Apache-2.0, 26,742 stars | Proprietary |
Grok Build fires 14 lifecycle events: SessionStart, SessionEnd, UserPromptSubmit, PreToolUse, PostToolUse, PostToolUseFailure, PermissionDenied, Stop, StopFailure, Notification, SubagentStart, SubagentStop, PreCompact and PostCompact. Only PreToolUse can block. Everything else is fail-open: a timeout, crash or malformed output is recorded in the session and the tool call proceeds anyway.
Claude Code documents 33 as of v2.1.271, a superset of Grok's plus Setup, InstructionsLoaded, UserPromptExpansion, MessageDisplay, PermissionRequest, PostToolBatch, TaskCreated, TaskCompleted, TeammateIdle, ConfigChange, CwdChanged, DirectoryAdded, FileChanged, WorktreeCreate, WorktreeRemove, PreModelSwitch, PostModelSwitch, Elicitation and ElicitationResult. The hooks system also supports prompt-based and agent-based handlers, not just shell commands and HTTP endpoints.
Grok Build: Apache-2.0 and ACP
The harness is on GitHub at 26,742 stars under Apache-2.0, so undocumented behavior is readable rather than guessed at. Full Agent Client Protocol support lets other apps drive the agent.
Claude Code: 33 Hook Events
33 lifecycle events including model-switch, worktree and file-change hooks, plus prompt-based and agent-based handlers. Two events can block a tool call rather than one.
A wire-level analysis of Grok Build 0.2.93 published in July 2026 found the CLI uploading the entire repository, including files the agent never read, as git bundles via POST /v1/storage to a Google Cloud Storage bucket named grok-code-session-traces. On a 12 GB repository the author measured 5.10 GiB moved on that channel against 192 KB on the model channel, and demonstrated recovering a never-read file verbatim from the captured bundle. The analysis reached the front page of Hacker News at 539 points. Its own update notes that xAI subsequently disabled the uploads server-side and added a /privacy command, which the author wire-tested as retention-only rather than transmission-blocking. The /privacy command is in the current Grok Build command table, described as "Show or toggle privacy and data-retention status." Run it before you point the agent at a repository you do not own.
When to Use Which
Choose Grok Build When
You Want to Read the Harness
Apache-2.0 on GitHub at 26,742 stars. Undocumented limits like the 32-subagent cap and the queue-versus-fail behavior are in the source, not in a support ticket.
Work That Decomposes Cleanly
Up to 32 concurrent subagents in general-purpose, explore and plan types. Migrating 40 files to a new API or sweeping a lint rule parallelizes; a coupled refactor does not.
Token Cost Matters More Than Context
grok-4.6 costs $2 in and $6 out per million tokens against Claude Opus 5's $5 and $25. grok-build-0.1 is cheaper still at $1 and $2 on a 256K context.
ACP Orchestration Matters to You
Grok Build speaks the Agent Client Protocol, so other apps can drive it as a backend. Claude Code's equivalent is its own subagent and Remote Control model.
Choose Claude Code When
Complex Multi-File Refactors
A 1M token context holds roughly 555,000 words, enough for most repositories at once. Single-agent coherence produces internally consistent changes across coupled modules.
You Gate the Agent With Hooks
33 lifecycle events against Grok Build's 14, with two blocking events instead of one, plus prompt-based and agent-based handlers and per-command allowed_domains in auto mode.
Verified Price at Entry
Claude Pro is $20 per month, or $17 billed annually, and includes Claude Code. Max starts at $100. xAI's consumer pricing page was not reachable for verification on September 14, 2026.
Published Head-to-Head Scores
73.4% on CursorBench 3.2 for Fable 5.1 and 70.0% for Opus 5, against 69.9% for Grok 4.6. It is the only benchmark both labs report on the same version.
| Priority | Best Choice | Why |
|---|---|---|
| Highest published coding score | Claude Code | 73.4% CursorBench 3.2 vs 69.9% for Grok 4.6 |
| Cheapest tokens | Grok Build | $2 in / $6 out vs $5 / $25 for Opus 5 |
| Auditable, forkable harness | Grok Build | Apache-2.0, 26,742 stars, no tags or releases |
| Large codebase reasoning | Claude Code | 1M token context vs 500K on grok-4.6 |
| Massive parallel fan-out | Grok Build | 32 concurrent subagents by default |
| Hook-gated automation | Claude Code | 33 lifecycle events vs 14, two of them blocking |
| Agent protocol (ACP) | Grok Build | Native ACP support for orchestration |
| Verified subscription price | Claude Code | Pro $20/mo, Max from $100/mo, published |
Frequently Asked Questions
What is Grok Build?
Grok Build is xAI's terminal coding agent. It shipped in early beta on May 14, 2026, opened to SuperGrok and X Premium+ subscribers on May 25, 2026, and became open source on July 14, 2026 at github.com/xai-org/grok-build under Apache-2.0. The CLI is at version 1.0.24 as of the September 9, 2026 monorepo sync. It runs grok-4.6 by default, cycles plan, auto and always-approve modes with Shift+Tab, spawns general-purpose, explore and plan subagents, and supports MCP servers, hooks, skills, plugins, worktrees, headless mode via -p, and ACP.
How does Grok Build's Arena Mode work?
It does not, any more. Arena Mode is no longer a documented Grok Build feature. May 2026 launch coverage described an automatic best-of-N evaluator that scored competing agent outputs. As of September 14, 2026 the string "arena" does not appear anywhere in the documentation at docs.x.ai, including its llms.txt index, and no /arena command exists in the command table. xAI has not published a note explaining the removal. The current equivalent is manual: spawn several subagents on the same task and compare their summaries yourself.
How much does Grok Build cost compared to Claude Code?
On the API, grok-4.6 costs $2 per million input tokens and $6 per million output tokens below a 200K prompt, with cached input at $0.50, doubling to $4 and $12 above 200K. Claude Opus 5 costs $5 in and $25 out, Claude Sonnet 5 $2 and $10, Claude Fable 5.1 $10 and $50. On subscriptions, Claude Code is included in Claude Pro at $20 per month, $17 billed annually, and in Claude Max from $100 per month. Grok Build signs in with a SuperGrok or X Premium+ subscription, but xAI's consumer pricing pages returned 403 and 404 when checked on September 14, 2026, so the widely repeated $299 per month SuperGrok Heavy figure could not be verified.
Which has better benchmark scores?
CursorBench 3.2 is the only benchmark both vendors publish. Anthropic reports 73.4% for Claude Fable 5.1 at max effort, 70.5% for Fable 5, 70.0% for Claude Opus 5 and 67.2% for GPT-5.6 Sol. xAI reports 69.9% for Grok 4.6. Anthropic has not published a SWE-bench Verified score for Opus 5; its last published figure was 80.84% for Claude Opus 4.6 in February 2026, which is why comparisons still quoting 80.8% as a current Claude Code score are stale. Terminal-Bench does not compare: Anthropic reports version 4.0, xAI reports version 3.0.
Does Grok Build support MCP servers and hooks?
Yes. Grok Build supports MCP servers, 14 hook lifecycle events, skills, plugins, marketplaces, worktrees, a four-profile sandbox, background tasks and ACP. It also reads Claude Code's files directly: AGENTS.md, CLAUDE.md, CLAUDE.local.md, .claude/rules/ and .cursor/rules/ all load as project rules, and hooks in .claude/settings.json are read with Claude tool names such as Bash, Read and Edit mapped onto Grok's automatically. Claude Code documents 33 hook events as of v2.1.271.
How do I use Grok in Claude Code?
Set ANTHROPIC_BASE_URL to https://api.x.ai/v1 and ANTHROPIC_AUTH_TOKEN to an xAI key, and Claude Code will talk to xAI's Anthropic-compatible /v1/messages endpoint. xAI's REST reference files that endpoint under "Legacy & Deprecated" and states the Anthropic SDK compatibility is fully deprecated, directing developers to the Responses API or gRPC, with no shutdown date published. Claude Code also degrades on a non-first-party host: MCP tool search is off unless ENABLE_TOOL_SEARCH=true, and Remote Control is disabled as of v2.1.196. Anthropic documents third-party routing only for Bedrock, Google Cloud, Microsoft Foundry and LLM gateways, all serving Anthropic models.
Related Comparisons
Faster Code Transformations for Any Agent
Morph Fast Apply merges LLM code edits at 10,500+ tokens/sec with 98% accuracy. Works with Grok Build, Claude Code, or any AI coding tool through the API.
