SpaceXAI launches Grok 4.7 for long-running coding work

SpaceXAI launches Grok 4.7 for long-running coding work
News

SpaceXAI released Grok 4.7 on 21 September 2026 as a new model for coding and knowledge work. The company says the model uses a larger base model and a longer reinforcement-learning run focused on difficult tasks that can take hours. It is designed to keep working through long jobs, check its own work more carefully and manage larger contexts than Grok 4.6.

The model is available immediately in Cursor and Grok Build, through the Grok API and through third-party coding harnesses, model routers and cloud platforms. GitHub also announced a gradual rollout in GitHub Copilot, where it is aimed at agentic coding and complex, multistep workflows. Copilot customers will be able to select it in Visual Studio Code, Visual Studio, the Copilot CLI, the cloud agent, JetBrains, Xcode and Eclipse, subject to plan and administrator settings.

SpaceXAI lists a 500,000-token context window for the API and supports low, medium, high and xhigh reasoning effort. Its starting price is 2 dollars per million input tokens and 6 dollars per million output tokens, the same list price shown for Grok 4.6. A faster variant is offered at twice those token rates. SpaceXAI reports improvements on CursorBench, DeepSWE, AA Briefcase, Terminal-Bench, legal work and electrical engineering. Those scores are provider-reported comparisons, so they should not be treated as independent proof that Grok 4.7 is best for every task.

Independent checks already show why the details matter. SiliconANGLE reported the release and its pricing, while security company XBOW tested an early version and found mixed results: Grok 4.7 was slightly weaker than Grok 4.6 in one fixed exploit-crafting setup but improved substantially when used with the Build orchestration environment. That suggests the surrounding agent harness can affect results as much as the model name.

The release matters because long-running agents are moving from demonstrations into common developer tools. A model that can handle larger context, use tools in smaller steps and verify intermediate work may reduce the amount of manual steering needed for codebases, documents and research tasks. It does not remove the need for review: higher token use, uneven benchmark results and gradual availability mean teams still need to measure cost, reliability and security in their own workflows before trusting it with consequential changes.

Source x.ai