⚡ Uncle Cat AI Radar
ModelsAgentsIndustry

xAI Releases Grok 4.7 for Coding and Knowledge Work

xAI says Grok 4.7 improves long-running coding and professional tasks while launching at $2 per million input tokens.

A larger model at a lower price

xAI has released Grok 4.7, describing it as the company’s strongest model so far for coding and knowledge work. The model uses a larger base model than Grok 4.6 and was trained with a longer reinforcement-learning run focused on tasks that require sustained reasoning over several hours.

The release is available through the Grok API, Grok Build, Cursor, third-party coding harnesses, model routers and cloud platforms. xAI lists pricing at $2 per million input tokens and $6 per million output tokens for prompts below 200,000 tokens. The API version has a 500,000-token context window and accepts text and image inputs, while producing text output.

Better on long tasks, but not universally ahead

xAI reports meaningful gains over Grok 4.6 on several evaluations. Its CursorBench 4.0 score rose from 40.4% to 46.3%, while its Terminal-Bench 4.0 score increased from 20.3% to 38%. Grok 4.7 also improved on legal work, clinical reasoning and electrical-engineering tests. On DeepSWE, it reached 71% using high effort.

The results do not put it first across the board. xAI’s own table shows the model behind Anthropic’s Fable 5.1 on some software and terminal tasks, while remaining competitive with other frontier systems. The company also presents a new safeguard stack, reporting lower success rates on risky cyber prompts and stronger performance on its biosafety test.

Why it matters

Grok 4.7 matters less as a simple leaderboard event than as a price-and-availability move. A frontier model aimed at multi-hour work is entering coding tools and agent platforms at prices that could encourage more persistent workloads. Whether the claimed gains survive independent testing, especially outside xAI’s selected benchmarks, will determine if this is a durable shift or an aggressive launch position.

Uncle Cat take

The important number is the 38% Terminal-Bench score: xAI is buying credibility in sustained tool use, but independent agent runs still decide whether the discount is real value.

Sources