Grok 4.6 Released: How the Model Matched GPT-5.6 Level at Half the Price
Anthropic
OpenAI
Moonshot AI
SpaceXAI released Grok 4.6, which according to the company matches the level of the most powerful models on the market at half the price: $2 per million input tokens and $6 per million output tokens. The model scores 61 on Artificial Analysis' Intelligence Index, equal to GPT-5.6 Sol, and has already become available in Grok Build, Cursor, Grok Bot, and via API.
SpaceXAI introduced Grok 4.6, claiming it reaches the level of the strongest models at half the cost of competitors: pricing is $2 per million input tokens and $6 per million output tokens. The model is already accessible in the Grok Build agent, in Cursor, in the Grok Bot, and through the API, with double usage limits for the first week in Cursor and Grok Build. The key figure is 61 points on the Intelligence Index of the independent analytics platform Artificial Analysis, which aggregates nine benchmarks. This matches GPT-5.6 Sol at maximum reasoning mode, is one point less than Claude Fable 5 Max, and is 5 points higher than Grok 4.5 from a month earlier. Analysts say SpaceXAI has returned to the frontier of intelligence, with only Anthropic ahead. On detailed benchmarks, Grok 4.6 leads on GDPVal-AA v2 for office knowledge work with 1753 points vs 1741 for Fable 5 and 1728 for Sol, and on AA-Briefcase with 1577 vs 1574 and 1502. It also tops the legal Harvey LAB with 15.8% compared to 11.3% for Fable 5 and 2.5% for Sol. However, it notably lags on Terminal-Bench v3.0: 26% vs 34.6% for Sol and 34.1% for Fable 5, but on the older v2.1 version it scores 88.4%, second only to Claude Opus 5. It also loses on some coding benchmarks: on DeepSWE it trails Sol (73%) and Fable 5 (70%), and on FrontierCode and APEX-SWE it loses to Fable 5. The most interesting aspect is how parity was achieved: according to Musk, Grok 4.6 is built on the same 1.5 trillion parameter base as Grok 4.5, so the model did not grow. Instead of scaling, the company ran another, longer training cycle on the same base with curated synthetic data on reasoning and engineering topics and an improved optimizer, then rebuilt the SFT and RL stages. SFT trajectories were regenerated by Grok 4.5 itself, filtering out problematic examples with model checks—effectively, the previous version trained the next one. In Cursor, now owned by SpaceXAI, they note that Grok 4.6 builds the structure and visual language of an app from the first pass based on an idea description, and on long trajectories it shows significantly more self-checking: it tests its own work before moving on. RL tasks during training ranged from optimizing computational kernels to web development and CAD. The economy remains from Grok 4.5, released July 8: the same $2/$6, which according to Artificial Analysis is over 60% lower than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). The Intelligence Index task costs $0.84 for the model, like Kimi K3, but with slightly more intelligence, placing it exactly on the Pareto frontier of intelligence versus price. Its signature token efficiency persists: on AA-Briefcase, Grok 4.6 completes a long agentic task in an average of 53 turns and half a billion input tokens, versus 103 turns and two billion for Claude Opus 5. Only cache hits increased, from $0.3 to $0.5 per million. SpaceXAI has a strong release pace: Grok 4.5 debuted July 8, Grok 4.6 August 12, and Grok 4.7 with 2.1 trillion parameters Musk estimated on the August call to be three to four weeks away, i.e., by late summer. Combined with the purchase of Cursor and the Grok Bot agent launched earlier, Grok is transforming from a meme into a product line with its own development environment, agent, and model pipeline.
- Abbreviations
- API = Application Programming Interface — программный интерфейс приложения
- SFT = Supervised Fine-Tuning — обучение с учителем
- RL = Reinforcement Learning — обучение с подкреплением
- CAD = Computer-Aided Design — система автоматизированного проектирования
Source: Habr — хаб ML —
original
