What happened?

On 21 September 2026 SpaceXAI shipped Grok 4.7. The news post and the API release notes carry the same day. The model id is grok-4.7. It is in Cursor, in Grok Build, and on the xAI API. Third-party harnesses and routers can serve it too. OpenRouter, Vercel, and Cloudflare are named on the docs page.

This is a coding and knowledge-work model, not a new chat personality. Context window is 500,000 tokens. Input is text and image. Output is text only, with no text output cap on the card. Knowledge cutoff is May 2026. Reasoning effort is low, medium, high (the default), or xhigh. On the Responses API it always returns encrypted reasoning, even if you did not ask for it. Pass those items back on the next turn or the model loses the thread.

SpaceXAI says 4.7 uses a larger base than Grok 4.6, with a longer reinforcement-learning run aimed at tasks that take many hours. It is better at checking its own work and at holding a long context. They also trained it to understand the Grok Bot harness, so the model and the always-on teammate are no longer strangers.

Pricing is where the two official pages disagree, and both should stay visible. The news post says it starts at $2 per million input tokens and $6 per million output tokens, served at the same price and speed as Grok 4.6, plus a fast variant with twice the output speed at twice the price. The API release notes are the sticker you actually bill: below 200,000 prompt tokens, $2 input / $0.50 cached input / $6 output per million. Above that, $4 / $1 / $12. The fast variant is the same model at 2x those rates, or 1.5x on long context. It is only in Cursor and Grok Build. It is not on the public API, and it is not in Grok Build's free tier. The US regional endpoint (us.api.x.ai) adds a 10% premium and keeps inference in the United States.

The headline on the news post says "twice as fast, at half the price of comparable models." The same page says 4.7 is served at the same price and speed as 4.6. Those are different comparisons. Half-price is versus other frontier stickers on their chart (GPT-5.6 Sol at $4 / $20, Fable 5.1 at $10 / $50). It is not a claim that 4.7 is twice as fast as 4.6.

Their own table, not an independent eval:

  • CursorBench 4.0: 46.3% for 4.7, 40.4% for 4.6, 41.7% for GPT-5.6 Sol, 51.8% for Fable 5.1.
  • DeepSWE v1.1: 71.0% for 4.7 at high effort, 65.2% for 4.6, 72.7% for GPT-5.6 Sol, 70.0% for Fable 5.1.
  • EEBench: 64.0%, against 53.0%, 39.4%, and 56.4%.
  • AA Briefcase v1.1: 1,657, against 1,546, 1,487, and 1,678.
  • Terminal-Bench 4.0: 37.6%, against 20.3%, 37.3%, and 57.9%.
  • Harvey Legal Agent Benchmark: 19.6%, against 15.8%, 2.5%, and 6.7%.
  • HealthBench Professional: 56.7%, against 48.5%, 60.5%, and 62.1%.
  • GDPval Elo: Fable 5.1 at 1,735, Grok 4.7 at 1,695, Grok 4.6 at 1,605, GPT-6 Astra at 1,542.

On the long coding bench it beats 4.6 and Sol and still trails Fable. On the legal-agent row the gap versus Sol and Fable is large. On clinical reasoning it improved on 4.6 and still sits under Sol and Fable. Read the row, not the headline.

Safety, in their words: a new safeguard stack, strongest refusal and jailbreak resistance they have tested, 62.4% on LatchBio's biosafety benchmark, and 3.3% of risky dual-use prompts let through on HackerBench v0.3. Select cybersecurity partners get invite-only red-team access. That is a lab claim, not a third-party audit.

Why this is interesting

  • The useful jump is the long job, not a new brand - Same sticker as 4.6 on the short prompts. The scores that moved are the ones that take hours: Terminal-Bench, CursorBench, briefcase work, documents.
  • Cached input is the real price - $0.50 per million cached tokens, if you stay under 200k and you actually hit the cache. The docs tell you to set a prompt cache key or you pay full input on a cold server.
  • Fast is a product, not an API switch - Twice the speed lives in Cursor and Grok Build. If you call grok-4.7 yourself, you get the standard tier.
  • It is wired for Grok Bot - Training on that harness is the bridge to the always-on teammate. The model card and the bot are one stack now.
  • The chart is theirs - Frontier, under Fable on the hardest coding row, ahead of 4.6 almost everywhere they printed. Treat it as a vendor table until someone else reruns it.

What it is not

Not twice as fast as Grok 4.6. Not a public fast tier. Not an independent leaderboard win over Fable 5.1. Not a new modality: still no native audio or video out. Not a reason to skip the cache key. The US endpoint is not the same price as the global one.

Bottom line

Grok 4.7 is the 4.6 successor you can call today, at the same short-prompt sticker, better on the long jobs SpaceXAI chose to print. Use the release-note prices, not the blog's opening line. If the work is a multi-hour coding or office loop, it is the Grok to try. If you needed the fast lane, that lane is inside Cursor and Grok Build, not on the raw API.