News

Grok 4

2 min read
Grok 4.6 reached the top of the price/performance list upon its release: the smart index is tied with GPT-5.6 Sol, but the price is more than 60% cheaper. 5 (62 points). This achievement puts Grok 4.6 into the first echelon of global cutting-edge models, competing head-to-head with the flagship products of OpenAI and Anthropic. The model extends training based on Grok 4.5, with core improvements including using curated data generated by the model for inference and advanced technical concept training, combined with high-quality engineering data and an improved optimizer. The key change in the training method is data closure: Grok 4.5 is used to regenerate supervised fine-tuning trajectories covering inference, agent tool usage, and areas such as STEM, software engineering, and knowledge work. Models are also trained on a wide range of agent reinforcement learning tasks, including knowledge work, general coding, kernel optimization, web development, and computer-aided design, a strategy clearly targeting core workloads of the "agent era." Agent benchmarks have soared across the board, and long-term tasks are surprisingly labor-saving. In the agent benchmark test, Grok 4.6 performed particularly well. GDPval-AA v2 reached 1753 Elo, second only to Claude Opus 5; DeepSWE v1.1 increased to 65.9%, a significant jump from Grok 4.5's 54%; APEX-Agents soared from 47.1% to 57.5%; Terminal-Bench v3.0 also increased from 15.7% to 26%. In addition, it achieved 69.9% on CursorBench v3.2 and 61.3% on FrontierCode v1.1, achieving significant breakthroughs in both coding and agent capabilities. Of greater concern are the cost-efficiency advantages. Grok 4.6 pricing remains at $2 per million input tokens and $6 per million output tokens, which is more than 60% lower than Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). In the AA-Briefcase long-term knowledge work benchmark, Grok 4.6 was completed in an average of just 53 rounds and about 500 million input tokens.