Grok 4.7 achieves improved coding benchmark scores at same price point
2 Sep 21 12:23 PM · 4d ago · 1 article · 1 post · 2 sources · development 2 of 2
Grok 4.7 maintains the $2/$6 token rates of version 4.6 while improving coding performance to 46.3% on CursorBench 4.0, up from 40.4%. A faster Cursor-only tier is available at double the standard rate.
“Grok 4.7 ships at the same $2 / $6 token price as 4.6, with a faster Cursor-only tier at double the rate. CursorBench 4.0: 46.3%, up from 40.4%”
gadgetry@techhub.socialSpaceXAI AI model developer
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
What was reported 1 claim about this development
-
first by Seeking Alpha, 4d ago
What people said 24 voices · best of 38 · verbatim
-
🚨 SpaceXAI just released Grok 4.7, and the most interesting number may not be a benchmark score. It’s the price. At $2 per million input tokens and $6 per million output tokens, Grok is pushing into the performance range of models that can cost several times more per task.
-
It's definitely gotten better at image->html workflows. Here's a test comparing Astra (currently SOTA at this) vs Grok 4.7:Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...Astra's build: https://html.non.io/annui/Grok's build: https://html.non.io/Annui-grok/Additional prompt instructions: "Add scrolling clouds behind the…
-
G
Grok 4.7 ships at the same $2 / $6 token price as 4.6, with a faster Cursor-only tier at double the rate. CursorBench 4.0: 46.3%, up from 40.4% Will you test Grok 4.7 in Cursor, or deploy it directly via custom API pipelines? https://www. developer-tech.com/news/spacex ai-grok-4-7-coding-reduced-token-cost/ # AI # Grok # Cursor # Developers # Tech
-
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is…
-
At every Grok release I commented on here how much I love Grok. It’s the consumer app and model I like the most, but my sentiment has changed. The usage limits on all the Grok subscriptions are now terrible, maybe because 4.6 and 4.7 (not in the app yet) eat more compote? I’m not sure what happenedMy SuperGrok subscription previously easily lasted…
-
https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level.Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
-
Initial impressions, Grok 4.6 for me just didn't really hack it for any usecase I tried. I seem to have a floor for my usecaseses (coding and a bunch of agentic workflows) and Sol/Opus are above some kind of intelligence floor.4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks…
-
It's a real dud IMO, worse than 4.6. I told it to fix a depth-testing issue with a WebGL scene and provided it a screenshot: told me it was fixed but it wasn't. I tell it to try again and it says it's "FOUND THE ROOT CAUSE!" then hasn't fixed it.I told it to compose an image (putting headgear on top of a head) - kept getting it completely wrong…
-
I've been using 4.6 for some one-off game mods/utilities and it has done very well. "I have a very niche keyboard (Moonlander) and I play this very niche space sim, make me a SVG keyboard cheatsheet for it". Told me to grab keymap.c for the keyboard and inputmap.xml for the game's key bindings, churned for a while, then spit out a pretty good…
-
> My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish.Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in…
-
I've put Grok 4.7 on the Redactle LLM benchmarks. It's a bit of a silly eval since it's a puzzle game but it tests omniscience really well.Grok 4.7 is near the top of the board. A significant improvement over Grok 4.6 but still not as good as Gemini 3.8 Flash which is very cheap and fast too.
-
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and…
-
Check out Grok 4.7 High ability at doing 3d scenes in threejs at threejseval.com https://threejseval.com/models/grok-4.7-highAlso go vote on https://threejseval.com so you can help evaluate how it performs compared to other models!
-
No doubt xAI has seen rapid progress, but it's been several months of them being "just behind" OpenAI and Anthropic. It seems the gap between just behind the frontier and pushing it is a lot wider than most people thought it was a year ago, and that's why a clear third contender in the frontier model space has yet to materialize.
-
as someone who is limited by amazon bedrock support at work (no idea why we got stuck with the worst one) - grok is literally the only budget-ish model option, so nice to see it updated, Sol and Opus are just too rich for my blood. Luna is good but so slow at getting things done (tps wise it's fast)
-
the AA numbers are generationally bad. double token use (the one thing Grok was good at was low reasoning usage!) to gain 5% in the benchmark score. with reportedly a larger model. maybe it shows gains IRL but wow, I've never seen a new generation model look so underwhelming compared to the last.
-
The CSAM generation model got an upgrade. The sad part about this is that I bet it still generates CSAM. Given that the owner of the company has made a nazi salute in public and thinks CSAM generating models are cool, I don't think they addressed the issue of this generating CSAM.
-
I don't fully understand what moves the needle further for these frontier models? Is the training data more valuable ? The training process ? The harness ?I know they are all important but where are they (all the frontier labs) really pushing to get incremental gains?
-
It’s a shame this model has such negative political baggage associated with it. It’s the only one I decided not to run in my LLM benchmarks[1].1 -
-
FYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.
-
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
-
> Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do.Lol, probably because Tesla's software stack is buildroot based. I'll bet that was in the training data.
-
> ClaudishI do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.
-
As anthropic/openai subscription allocations get squeezed you'll see more people using "second rate" closed models like grok. The token allowance with a Cursor subscription is crazy.
All 2 developments of SpaceXAI releases Grok 4.7 with coding improvements… →
Hacker NewsNewswiresMastodonX