Grok 4.7 fails most long-running tasks despite design for hours of work
xAI's latest agent struggles with stamina despite being built to handle multi-hour coding tasks.
What to know
- xAI released Grok 4.7 specifically designed to autonomously work for hours on coding tasks, making dozens of decisions without human intervention.
- Early reports show Grok 4.7 fails most of the time, significantly underperforming its stated capability for sustained task execution.
- The release has drawn immediate skepticism from observers on social media.
“A coding agent running for hours can make dozens of decisions as it edits files, runs tests, and works through”
Amanda Caswell, The New Stack writer · The New Stack ↗ · Sep 21, 3:13 PM
xAI Developer of Grok 4.7
How it unfolded 1 development · click the chart to see its coverage articlesposts
-
1
@synthwavedd calls Grok 4.7 launch disappointing
An X user responds to the launch with a negative assessment, saying the model is 'pretty terrible' and wishes better luck next time.
“Well, they launched it, and as expected... Grok 4.7 is pretty terrible. Better luck next time…”
— @synthwavedd -
Well, they launched it, and as expected... Grok 4.7 is pretty terrible. Better luck next time
-
-
background
Grok 4.7 fails most of the time despite long-duration design — The New Stack reports that Grok 4.7 fails most of the time, undermining its core premise of sustained autonomous work over hours.
-
background
xAI releases Grok 4.7 agent designed for multi-hour autonomous work — xAI launched Grok 4.7, a coding agent built to operate for hours while making dozens of decisions as it edits files, runs tests, and works through tasks.