SpaceXAI has launched Grok 4.7, its latest frontier AI model, with a bigger underlying model, longer reinforcement-learning training and a particular focus on coding and complex knowledge work.
The company is positioning Grok 4.7 as a substantial upgrade over Grok 4.6 rather than a minor iteration. According to SpaceXAI, the new model has been trained on a more difficult mixture of tasks, including problems designed to take hours to complete, while improvements to context management and self-verification are intended to make it more dependable during longer workflows.
Coding is predictably one of the headline battlegrounds. SpaceXAI reports a 46.3% score for Grok 4.7 xHigh on CursorBench 4.0, compared with 40.4% for Grok 4.6 High. It also claims 71% on DeepSWE v1.1 at high effort and 38% on Terminal-Bench 4.0.
Those figures put Grok 4.7 into direct competition with other frontier models, although company-published benchmark comparisons deserve the usual caution. Different reasoning settings, test configurations and costs can materially affect results, so benchmark leadership doesn’t automatically translate into a better model for every real-world workload.
SpaceXAI is also pushing beyond software development. Its published results show improvements in electrical engineering, multi-hour office work and legal tasks, while the company says Grok 4.7 has become better at generating documents and presentations. The model has additionally been trained to understand the company’s Grok Bot harness natively, targeting conversational and broader knowledge-work applications.
Safety is another area receiving attention. SpaceXAI says Grok 4.7 uses an entirely new safeguard system and reports stronger jailbreak resistance than its previous models. The company claims the model allowed 3.3% of risky dual-use prompts through on its HackerBench v0.3 evaluation, while attempting to maintain usefulness for legitimate cybersecurity work. Select security partners are also receiving invite-only access to red-team capabilities.
Perhaps the more immediately competitive number is price. Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens. A faster version offering twice the output speed costs twice as much.
Grok 4.7 is available through the Grok API, Grok Build, Cursor, third-party coding tools, model routers and cloud platforms.
The bigger test now won’t be another benchmark chart. It will be whether Grok 4.7’s claimed gains in long-running, autonomous work survive contact with the messy projects developers and professional users actually give it.
