xAI Launches Grok 4.7 for Coding and Knowledge Work — Stronger Than 4.6, Still Mid-Pack on Independent Scores
Islamabad / Copenhagen — 22 Sep 2026 (PKT). xAI has released Grok 4.7, calling it the company’s most capable Grok yet for coding and knowledge work. The model landed around 21 September 2026 and is already showing up in developer tools — even as independent scorecards place it in the middle of the frontier pack rather than at the very top.
What xAI says changed
According to xAI’s launch note, Grok 4.7 sits on a larger base model than Grok 4.6. The company says it ran longer reinforcement learning — a training method that rewards the model for solving harder problems — on tasks that can take many hours. xAI also claims better self-verification (the model checking its own answers), stronger long-context handling, and native training on the Grok Bot harness used for conversational and agent-style work.
A new safeguard stack is part of the pitch. xAI describes it as its best-calibrated safety system to date. Those are company claims; outside labs have not yet published a full independent safety audit of the stack.
Price and where you can use it
Pricing matches Grok 4.6: starting at $2 per million input tokens and $6 per million output tokens. xAI also offers a fast variant that doubles output speed at double the price. Availability, per the company, includes Cursor, Grok Build, the Grok API, third-party coding harnesses, and cloud platforms.
Separately, GitHub’s 21 September 2026 changelog said Grok 4.7 is rolling out in GitHub Copilot for Pro, Pro+, Max, Business, and Enterprise plans across VS Code, Visual Studio, JetBrains and other Copilot surfaces. The rollout is gradual, GitHub noted.
Company benchmarks — attribute to xAI
xAI published its own comparison table. On CursorBench 4.0, which stresses longer coding tasks, Grok 4.7 scored 46.3% versus 40.4% for Grok 4.6. On Terminal-Bench 4.0, the jump was sharper: 38.0% against 20.3%. On EEBench, an electrical-engineering test, xAI reported 64%. The company also showed gains on multi-hour office work and several professional-domain quizzes.
Those numbers come from the vendor. They help explain the “better than 4.6” message — especially price-performance on CursorBench — but they are not a substitute for third-party evaluation.
Independent coverage: mid-pack, not the ceiling
Balanced reporting matters here. Artificial Analysis scored Grok 4.7 at 46 on its Intelligence Index, up from Grok 4.6, while still trailing leaders such as Claude Fable 5.1 and GPT-6 (reported around 53 on the same index by The Decoder). Decrypt and The Decoder both framed the release as a solid step up that remains behind Anthropic and OpenAI models on several public indices.
In plain terms: Grok 4.7 looks more useful for developers who already live in Cursor, Copilot, or the Grok API — especially if cost and speed matter as much as raw leaderboard rank. It is not, on today’s independent charts, the undisputed frontier leader.
Bottom line
For readers watching the 2026 AI tooling race, Grok 4.7 is a same-price upgrade with clearer coding-agent intent and a wider distribution footprint. Trust the company’s own benches for the “better than 4.6” story. Trust independent indices for where it sits against Claude Fable 5.1 and GPT-6. Both stories can be true at once.






















