Grok 4.6 arrives with a 500K context window and a push into long-running agent work

SpaceXAI released Grok 4.6 on August 12, the latest iteration of its flagship model family, aimed squarely at software development and long-running AI agents. The company, formed after SpaceX absorbed xAI in June 2026, positions the release as a response to growing demand for models that can sustain complex multi-step tasks: researching a subject, moving through a large codebase, or turning a rough concept into a working application.

The model keeps the 500,000-token context window of its predecessor, Grok 4.5, with a knowledge cutoff of February 1, 2026. It accepts text and images and produces text, with configurable reasoning levels and tool use spanning function calling, code execution, web search and X search. Pricing holds at $2 for each million input tokens and $6 for each million output tokens, with a faster variant at double the price.

SpaceXAI says Grok 4.6 ties GPT-5.6 Sol on Artificial Analysis’s composite intelligence index, with both scoring 61 against 62 for Claude Fable 5 Max. Independent testing by Artificial Analysis reaches a similar verdict: an Elo of 1753 on its GDPval-AA v2 gauge of practical agentic knowledge work, second only to Claude Opus 5, and close enough to Claude Fable 5 and to Qwen3.8 Max to fall within the margin of error. The same analysis found Grok 4.6 finishing tasks in about 53 turns and roughly 0.5 billion input tokens on average, compared with about 103 turns and 2 billion tokens for Claude Opus 5 at maximum effort, which gives it a large cost advantage on long agentic runs.

The vendor’s own benchmark table shows gains over Grok 4.5 everywhere, including an 11.9-point jump on DeepSWE and 10.3 points on Terminal-Bench, and first-place finishes on GDPVal-AA v2, on AA-Briefcase and on the Harvey LAB legal benchmark. On coding-specific measures the picture is more mixed: GPT-5.6 Sol Max still leads DeepSWE v1.1 with 73 percent and Terminal-Bench v3.0 with 34.6 percent.

Support journalism that values evidence, context, and accuracy above everything else.

Support 1ban.news

Training followed the pattern set by Grok 4.5 but went further: a longer supplemental run with curated model-generated data, an improved optimizer, and SFT trajectories regenerated by Grok 4.5 itself across reasoning, STEM, software engineering, agent harnesses and knowledge work, with problematic traces filtered out. Reinforcement learning covered agentic tasks in kernel optimization, web development and computer-aided design. The company says the model’s safeguards were calibrated alongside the expanded capabilities and tested in its widest-ever pre-deployment suite.

Availability mirrors the Grok 4.5 rollout: the model is live in Cursor and Grok Build, through the SpaceXAI API and via partners such as OpenRouter, Vercel and Cloudflare, with double included usage in Cursor and Grok Build during the launch week. The timing fits SpaceXAI’s rapid post-IPO cadence; the company, publicly traded as SPCX, has made coding and agentic work its competitive lane, and its acquisition of the Cursor editor gave it direct access to the developer market and real coding sessions for training data.

Sources: Introducing Grok 4.6 (SpaceXAI, Aug 2026); Introducing Grok 4.6 (Cursor, Aug 2026); Grok 4.6 (xAI Docs, Aug 2026); Grok 4.6 benchmarks and analysis (Artificial Analysis, Aug 2026); SpaceXAI releases Grok 4.6 (9to5Mac, Aug 2026)

Scroll to Top