
Grok 4.5: Coding Benchmarks, 80 TPS, and Token Efficiency
SpaceXAI's July 16 model posts 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-bench Pro while averaging 15,954 output tokens per task.
Read MoreZharfAI Research
Model release desk

SpaceXAI released Grok 4.6 on August 12, 2026, less than a month after Grok 4.5. The official announcement positions it as a frontier model for long-running agents, coding, research, and ambitious interactive or visual work. The accompanying 36-page model card adds capability and safety results that make the release more mixed than the launch table alone suggests. Meanwhile, Eric Zakariasson's field guide supplies the more useful operating lesson: a capable agent still needs an explicit definition of “done” and a way to inspect what it produced.
That distinction matters. Benchmark gains measure performance in named harnesses. A firsthand product essay shows how one experienced user worked with the model. Neither proves that Grok 4.6 will complete an organization's own jobs safely or economically. The practical question is whether its speed, judgment, and self-verification reduce the total effort required to reach an accepted result.

First-party media from SpaceXAI's Grok 4.6 release. It establishes the product identity; it is not independent evidence of capability.
SpaceXAI describes 4.6 as a continuation of Grok 4.5, not a new product category. The supplemental training run was longer and used curated model-generated reasoning data, advanced technical material, engineering data, an updated optimizer, and a revised training recipe. Grok 4.5 then regenerated supervised fine-tuning trajectories across several reasoning efforts, agent harnesses, and domains. Model-based checks filtered problematic traces before reinforcement learning.
The reinforcement-learning environments extended beyond general software work into knowledge tasks, kernel optimization, web development, and computer-aided design. SpaceXAI says this produced stronger first passes on interactive and visual applications, better persistence across many steps, and more instances of the model testing its own work before continuing.
Those are vendor-reported training and behavior claims. The release does not publish weights, the composition of its training corpus, or a reproducible recipe. Its model card does provide substantially more evaluation detail than the launch page, but it remains a first-party disclosure rather than an independent audit.
The model card describes Grok 4.6 as the latest release in a “1.5T-scale model family,” developed with Cursor. It says supplemental training used anonymized Cursor workflow data in addition to public, internally generated, and licensed material. This helps explain the emphasis on repository-scale work, but neither the data mix nor a reproducible training recipe is disclosed.
The card also uses a January 2026 pretraining cutoff, while the developer guide lists a February 1, 2026 knowledge cutoff. Those labels can describe different stages of a model pipeline; users should preserve both claims instead of turning them into a promise of current knowledge.
Most importantly, the card shows that performance and safety did not move in one direction:
Other results add context. Child-safety compliance stayed at 0%, and the model reached 100% refusal accuracy on the card's internal biological and chemical weapons suites. The card says biological and chemical capabilities remain below SpaceXAI's Frontier AI Framework thresholds. Yet ProtocolQA accuracy fell from 87.0% to 79.6%, and most bio-capability tests were intentionally run without safeguards to measure the underlying model. Cyber capability rose modestly on CyberGym and CVE-Bench; SpaceXAI says unrestricted testing by third parties validated that direction, but the card does not name the evaluators or publish their reports.
The most concrete agentic case study is also internal. An earlier checkpoint spent five hours testing 297 candidate inference optimizations, opened seven pull requests, and had three changes reach Grok Chat production. SpaceXAI attributes a combined 1.5% decode-throughput and 3.1% prefill-throughput gain to those changes. That is useful evidence of an observation-and-repair loop in a real engineering environment, but it is still a vendor-selected result on proprietary infrastructure.
The card therefore improves transparency without certifying the model. Several evaluations are internal implementations, competitor numbers often come from model cards or public leaderboards, and sample counts or run-to-run variance are not reported for many suites. SpaceXAI explicitly says Grok 4.6 is not intended to make autonomous high-stakes decisions in medicine, law, finance, or other safety-critical domains without oversight and domain validation.
The official table reports improvements over Grok 4.5 High across every listed evaluation. Grok 4.6 also matches GPT-5.6 Sol Max on the nine-benchmark Artificial Analysis Intelligence Index, while Fable 5 Max is one point higher. On individual tasks, the ordering changes.
| Evaluation | Grok 4.6 High | Grok 4.5 High | What it samples |
|---|---|---|---|
| AA Intelligence Index | 61 | 56 | Composite of nine evaluations |
| GDPVal-AA v2 | 1753 | 1526 | Professional knowledge work |
| CursorBench v3.2 | 69.9% | 66.7% | Agentic coding inside Cursor's harness |
| DeepSWE v1.1 | 65.9% | 54.0% | Repository-level software work |
| FrontierCode v1.1 Extended | 61.3% | 56.6% | Longer coding problems |
| APEX-Agents | 57.5% | 47.1% | Cross-domain professional agent tasks |
| Terminal-Bench v3.0 | 26.0% | 15.7% | Terminal task completion |
| AA-Briefcase | 1577 | 1313 | Work-product quality |
The largest relative jump in this selection is Terminal-Bench, but Grok 4.6 remains behind both cited comparison models there. It leads the displayed set on GDPVal-AA v2, APEX-Agents, AA-Briefcase, and Harvey LAB under the configurations shown. Fable leads CursorBench, DeepSWE, FrontierCode, APEX-SWE, and Terminal-Bench. A composite tie therefore does not mean interchangeable behavior.
Competitor figures on the launch page are the best self-reported or publicly available results, not one fully controlled run. Preserve model version, reasoning effort, scaffold, tools, time budget, and evaluation date before repeating any ranking. For a deployment decision, a private task set and accepted-output cost matter more than a headline average.
Zakariasson says he used Grok 4.6 for several weeks as a daily driver across coding and knowledge work. His examples include navigating provider consoles, functional and visual QA, inbox triage, an Excalidraw feature, a spreadsheet application, a strategy game, a board deck, 3D work, and Remotion videos. He found the combination of speed and capability important enough to move from asynchronous delegation toward shorter, synchronous loops.
This is unusually detailed practitioner evidence, but it is not an independent review. Zakariasson works with Cursor, Grok 4.6 launched in Cursor, and his post amplifies the official release. His comparisons are still useful because he says the paired projects used the same prompts in isolated workspaces. They should be read as a set of hypotheses to reproduce, not a purchasing verdict.
The strongest hypothesis is not “4.6 can build anything.” It is that the model's first-pass taste and speed make a shorter feedback loop productive. That can reduce specification overhead when a knowledgeable operator is present. It does not remove the operator.
The field guide compares a two-page spreadsheet specification with a three-sentence version. Zakariasson reports that the resulting applications were nearly identical. In his testing, ornamental phrases urging the model to work harder contributed little. Prompt length mattered because it changed who made the decisions: a long specification constrained the result; a short request delegated more design judgment to the model.
That is a better framework than “short prompts win.” Use a short brief when the solution space is open, the output is cheap to inspect, and the reviewer is ready to steer. Use a detailed specification when requirements are contractual, the domain has hidden constraints, or a missed condition is expensive. A good model can infer a plausible interface; it cannot infer an organization's unstated policy with authority.
Keep preferences separate from acceptance criteria. “Make it polished” describes direction. “A new user can import this fixture, edit a formula, undo the change, export the workbook, and obtain the expected file” defines evidence. Our software-engineering agent guide explains how to connect that evidence to tests, regressions, and reviewer effort.
In the spreadsheet comparison, the material improvement came after the prompt required the agent to verify function and design, then keep iterating until the application was production-ready. Zakariasson says that instruction caused the model to open the application, follow real user paths, test nested formulas, and repair failures. In another example, asking for a captured frame and a concrete defect list worked better than a vague request to improve 3D textures.
The important part is not a magic sentence. It is the availability of an observation-and-repair loop:
A summary that says the task is complete is not completion evidence. For code, require the running application and relevant tests. For research, require source-backed claims and contradiction checks. For a spreadsheet, recalculate known fixtures. For an external action, verify the exact destination and result without automatically repeating a side effect. The computer-use evaluation guide provides a fuller protocol for state, recovery, and action safety.
Web applications expose a machine-readable DOM, logs, routes, and screenshots, so an agent can inspect both structure and appearance. Zakariasson's field guide identifies a useful gradient: 3D adds spatial relationships that source code alone cannot confirm; video adds time; physics requires a sequence of states and causal behavior. One attractive frame cannot prove pacing, collision behavior, continuity, or interaction quality.
This explains why better visual taste does not eliminate steering. Give the model the observation instrument appropriate to the medium: multi-angle captures for 3D, sampled frames and full playback for video, trajectory logs for physics, and scenario traces for stateful products. Where the model cannot see the relevant dimension, assign that check to a human or another deterministic system.
Evaluation must also separate generation quality from verification quality. A model may produce a stronger first pass yet miss a broken interaction. Score both. A reliable workflow should prefer a slightly less impressive artifact that passes defined checks over a beautiful demo with an untested core path.
The official developer guide names the API model grok-4.6. It accepts text and image input and returns text, has a 500,000-token context window, and lists a February 1, 2026 knowledge cutoff. Reasoning effort can be set to low, medium, high, or xhigh. The Responses and Chat Completions APIs support function calling, web search, X search, and code execution.
Standard API pricing is $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. The launch page says a faster variant costs twice as much. The model detail page warns that requests beyond 200,000 context tokens use higher-context pricing, so the headline input rate should not be applied blindly to very large jobs.
For long sessions, SpaceXAI recommends a conversation-level prompt-cache key and context compaction. Track cache hit rate, compaction behavior, tool charges, retries, and container time; token list price alone does not describe agent cost.
At launch, Grok 4.6 is available in Cursor on all plans, is the default in Grok Build, and is offered through the SpaceXAI API, OpenRouter, Vercel, Cloudflare, Snowflake, and Databricks Mosaic AI. The model card says consumer availability on the web, mobile apps, and Grok in X is planned for a later date. Do not infer current consumer access from a general “try Grok” button.
Evaluate 4.6 beside the current production model and Grok 4.5 on the same immutable task snapshots. Include short fixes, unfamiliar repository work, research with conflicting sources, one office artifact, and one interactive or visual project. Give each model the same tools, permissions, wall-clock ceiling, and action policy.
For every task, run two prompt conditions: a detailed specification, and a short brief with the same non-negotiable acceptance criteria. Require the agent to show the running result and its verification evidence. Have reviewers score correctness, completeness, maintainability, visual judgment, unsupported claims, unsafe actions, and honesty about unfinished work without knowing which model produced the artifact.
Measure pass rate, first-pass acceptance, regressions, elapsed time, input and output tokens, cache savings, tool cost, reviewer minutes, and cost per accepted task. Interrupt some runs, change a dependency, expire a session, and inject misleading content into a document or issue. A long-running agent must recover safely, not merely continue for a long time.
Grok 4.6 is a credible step up from 4.5 across the launch suite and a more interesting release than a single composite score suggests. Its value proposition is the combination of frontier-level results, fast interactive use, stronger first passes, long-task persistence, and a comparatively low standard API price. The model card makes the decision less comfortable in a useful way: search, coding, professional work, and several refusal measures improve, while hallucination, self-harm handling, dishonesty under pressure, and sycophancy regress on the reported tests.
The field guide sharpens the deployment lesson: prompt cleverness is secondary to a clear finish line and an observation loop. Short prompts can be effective when the model's taste is good and inspection is cheap. For consequential work, write the requirements down, expose the real output to the agent, preserve evidence, and keep a human responsible for dimensions the model cannot reliably observe.

SpaceXAI's July 16 model posts 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-bench Pro while averaging 15,954 output tokens per task.
Read More
Z.ai's June 16 MIT release improves terminal, repository, tool-use, and long-running tasks through shared sparse indexing and controllable effort.
Read More
Moonshot's June 12 model lifts coding and MCP tool scores over K2.6 while cutting reasoning-token use about 30 percent, with full benchmark footnotes.
Read MoreIf this note maps to a real system in your organization, start with the services page or a shipped case study.