conv.

All stories
AIQuiet 5d · day 5

xAI's Grok 4.7 launches with low prices but trails Claude and GPT-6 in benchmarks

Elon Musk's latest model scores mid-pack on independent tests despite larger size and reinforcement learning improvements.

What to know

  • Grok 4.7 launches at $2 per million input tokens, undercutting Claude and GPT-6, but scores 46 on the Artificial Analysis Intelligence Index versus 53 for both competitors.
  • The performance gap widens on coding tasks: Grok 4.7 achieves only 26% on Terminal-Bench 4.0 versus 60% for GPT-6 Astra and 55% for Claude Fable 5.1.
  • The model includes 2.1 trillion parameters (40% larger than Grok 4.6) and training data from SpaceX, but follows xAI's consistent pattern of trailing frontier models despite significant resources.
  • Musk delayed the release at least five times since late July, ultimately launching with no waitlist through the Grok app, Cursor, Grok Build, and the xAI API.

“xAI has consistently undercut Anthropic and OpenAI on cost per token even as it trails them on raw capability, betting that "good enough, cheap, and everywhere" beats "best, but pricier" for the bulk of everyday use.”

Decrypt, Technology reporter · Decrypt ↗ · Sep 20

Elon MuskElon Musk xAI founderxAI Model developerClaude Fable 5.1 (Anthropic) Competing modelGPT-6 (OpenAI) Competing model

xAI's Grok 4.7 launches with low prices but trails Claude and GPT-6 in benchmarks
decrypt.co

How it unfolded 1 development · click the chart to see its coverage articles

Peak 3 pieces in two hours at Sep 21, 11 AM; 3 pieces over 5 days (3 articles) Sep 21, 11 AM — 3 pieces · 3 articles — Newswires 3Sep 21, 1 PM — quietSep 21, 3 PM — quietSep 21, 5 PM — quietSep 21, 7 PM — quietSep 21, 9 PM — quietSep 21, 11 PM — quietSep 22, 1 AM — quietSep 22, 3 AM — quietSep 22, 5 AM — quietSep 22, 7 AM — quietSep 22, 9 AM — quietSep 22, 11 AM — quietSep 22, 1 PM — quietSep 22, 3 PM — quietSep 22, 5 PM — quietSep 22, 7 PM — quietSep 22, 9 PM — quietSep 22, 11 PM — quietSep 23, 1 AM — quietSep 23, 3 AM — quietSep 23, 5 AM — quietSep 23, 7 AM — quietSep 23, 9 AM — quietSep 23, 11 AM — quietSep 23, 1 PM — quietSep 23, 3 PM — quietSep 23, 5 PM — quietSep 23, 7 PM — quietSep 23, 9 PM — quietSep 23, 11 PM — quietSep 24, 1 AM — quietSep 24, 3 AM — quietSep 24, 5 AM — quietSep 24, 7 AM — quietSep 24, 9 AM — quietSep 24, 11 AM — quietSep 24, 1 PM — quietSep 24, 3 PM — quietSep 24, 5 PM — quietSep 24, 7 PM — quietSep 24, 9 PM — quietSep 24, 11 PM — quietYesterday, 1 AM — quietYesterday, 3 AM — quietYesterday, 5 AM — quietYesterday, 7 AM — quietYesterday, 9 AM — quietYesterday, 11 AM — quietYesterday, 1 PM — quietYesterday, 3 PM — quietYesterday, 5 PM — quietYesterday, 7 PM — quietYesterday, 9 PM — quietYesterday, 11 PM — quietToday, 1 AM — quietToday, 3 AM — quietToday, 5 AM — quietToday, 7 AM — quietToday, 9 AM — quiet 1
Sep 22Sep 23Sep 24yesterdaynow · 10:09 AM ET
  1. 1

    Independent benchmarks show Grok 4.7 significantly trailing frontier models

    On the Artificial Analysis Intelligence Index (v4.3.2), Grok 4.7 scores 46 and lands mid-pack, versus 53 each for Claude Fable 5.1 and GPT-6. The gap widens on agentic coding: on Terminal-Bench 4.0, Grok 4.7 hits just 26 percent versus 60 percent for GPT-6 Astra and 55 percent for Claude Fable 5.1, even trailing DeepSeek V4.1 Flash at 27 percent.

    “a strong combination of intelligence, speed & low cost…”
    — Elon Musk, xAI founder · source
  2. background

    xAI releases Grok 4.7 with increased parameters and lower pricing — xAI launches Grok 4.7, featuring 2.1 trillion parameters (40% larger than Grok 4.6's 1.5 trillion), priced at $2 per million input tokens and $6 per million output tokens. The company says it spent longer on hard problems, double-checks answers more often, and includes supplemental training data from SpaceX including Starlink telemetry, manufacturing records, and engineering failure logs.

  3. background

    Musk first signals Grok 4.7 timeline, beginning pattern of delays — Elon Musk begins announcing the Grok 4.7 release, initially saying it would be ready in "four weeks out." Over the following two and a half months, he repeatedly pushes back the timeline through at least five iterations: "a few weeks," "3 to 4 weeks," "10 days" on September 1, and "needs a few more days to cook" on September 11.