xAI releases Grok 4.6 for long-running agents and visual work
xAI released Grok 4.6 on August 12, positioning the model as an upgrade for long-running agents and more ambitious interactive and visual work. It builds on Grok 4.5 and is designed to remain engaged across many steps, including researching an unfamiliar topic, analyzing information, working across a codebase and turning an idea into a functioning application. The model is available in Cursor and Grok Build, through the xAI API, and from partners including OpenRouter, Vercel and Cloudflare.
The training changes are aimed at sustained work rather than a single impressive answer. According to xAI, Grok 4.6 received a longer supplemental training run with curated reasoning and engineering data plus an updated optimizer and training recipe. It used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning settings and domains such as engineering and knowledge work. Model-based checks filtered problematic traces before reinforcement learning on a broad set of agent tasks.
xAI says Grok 4.6 reaches frontier performance across agentic coding and knowledge-work evaluations. On the Artificial Analysis Intelligence Index, a composite of nine benchmarks, the company reports that it matches OpenAI’s GPT-5.6 Sol. Longer project tests reportedly showed more self-testing and stronger first versions of visual work. These are developer-reported results, and the comparison figures come from published system cards or benchmark leaderboards. Independent testing of reliability over long sessions is still needed.
For developers and companies, the practical promise is continuity. A model that can preserve context, test its own output and improve a project through several rounds may reduce the supervision required for coding, research and document production. It also raises the cost of mistakes: an agent can repeat a faulty assumption across many steps or make changes throughout a codebase. Human review, restricted permissions, version control and clear stopping conditions remain necessary when the model can act rather than only advise.
Pricing starts at $2 per million input tokens and $6 per million output tokens; a faster variant costs twice as much. The release matters less as another benchmark update than as evidence of where model competition is moving: from short answers toward agents that can stay with complex work, verify progress and deliver usable artifacts. The real measure will be whether Grok 4.6 does that predictably outside xAI’s demonstrations.