Part of Apple's AI-Optimized Silicon Push · 2 stories · since Sep 15 · newest 1d ago
M5 Ultra Mac Studio reviews call it the best machine yet for local AIMKBHD says the machine now targets AI users over creators
6 Sep 21 10:06 AM · 2d ago · 7 articles · 10 posts · 20 comments · 5 sources · development 6 of 6
Reviewer Marques Brownlee posts that after a week of testing, the M5 Ultra Mac Studio's improvements now seem aimed at local AI workloads rather than video creators like himself.
“At this point, the improvements are not even targeted at me (a video creator) anymore. They're targeted at being the best machines for local AI work.”
@MKBHDFederico Viticci MacStories editorApple ManufacturerAntonio G. Di Benedetto The Verge reviewerMarques Brownlee (MKBHD) Tech reviewer/YouTuberRoman Loyola Macworld reviewer
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 4 claims were published in its stretch
-
first by iClarified, 2d ago · also PetaPixel
1 more headline
- Mac Studio With M5 Max Review: The Best Mac for Photo and Video We’ve Tested PetaPixel · 1d ago
-
first by Macworld, 2d ago · also TechRepublic
1 more headline
- Mac Studio M5 Max vs M5 Ultra: Is the $3,000 Upgrade Worth It? TechRepublic · 1d ago
-
first by CNET, 1d ago
-
first by The Verge, 1d ago
What people said 24 voices · best of 26 · verbatim
-
Been using the M5 Ultra Mac Studio for the past week and predictably it is the most powerful computer I've ever used. At this point, the improvements are not even targeted at me (a video creator) anymore. They're targeted at being the best machines for local AI work. There's
-
The Cost of a 27- inch fully loaded iMac from Apple $3,700 in 2011 is worth $5,510.05 today (fully loaded) used for 10 years.The Mac Pro Tower plus the Cinema Display monitor at the time was even more, similar to the Studio M5 ULtra today than the iMac which I believe was an upper middle computer?$5,700 in 2011 is worth $8,488.46 today$7,700 in…
-
P
Apple's 2026 Mac Studio with the M5 Ultra pushes performance (and pricing) to the limit. The new quad-die chip is a serious powerhouse for top-tier local AI inference, heavy software compilation, and high-throughput video rendering. Here's our review: https://www. pcmag.com/reviews/apple-mac-st udio-2026-m5-ultra
-
This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.It's already clear to me that M5 Mac Studio is…
-
Federico Viticci has an exhaustive review of the new Mac Studio with M5 and 256 GB of RAM. Lots of cool charts on performance. It sounds like a great Mac. The problem is that config will set you back $9k! For me, a $100/month subscription for frontier AI over many years is still better.
-
> "It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can…
-
N
M5 Ultra Mac Studio Review: https://www. macstories.net/stories/m5-ultr a-mac-studio-review-the-dream-mac-for-local-ai-agents/ Discussion: http:// news.ycombinator.com/item?id=4 9787313
-
A basic llama-bench on Qwen 3.8 27B UD-Q4_K_M gives pp512 3920 tok/s / tg128 81 tok/s on a 500W RTX PRO 6000 (should be similar speeds to a 5090, chip is basically the same, just less VRAM). With MTP3 this is 140 tok/s on mtp-bench.This is with llama.cpp. You can of course use vLLM/SGLang well on these cards and they're even faster. On vLLM w/…
-
J
https://www. macstories.net/stories/m5-ultr a-mac-studio-review-the-dream-mac-for-local-ai-agents/ This was pretty interesting!
-
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090: Qwen3.8 27B tokens/sec generation speed Prompt size 8K 64K 128K 256K RTX 5090 PC 59 51 44 n/a M5 Ultra 48 39 32 24 M3 Ultra 31 23.5 20 15 A whole bunch more comparison numbers in this section:
-
I really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in.These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over…
-
My mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am…
-
On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add https://huggingface.co/collections/z-lab/dflash-2 to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.
-
While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.This is a great article and bodes well for the M5, but we…
-
Those are some incredible graphs, that leap in prompt processing going from M3 to M5.Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think)./meta Here's a CSS filter that stops those nuisance chart animations, macstories.net##*:style(animation: none !important; transition: none !important)
-
Once you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB/s. IMO best ROI for most people is 'cheapest used Ultra with enough RAM'I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
-
I find it funny that the thing always mentioned with this machine is local AI. If you're a local model enthusiast, then maybe that makes sense, but I just don't see the economics working there.I ordered the same one for work so I could run more local agents at once (many iOS simulators and Xcode build processes).
-
When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
-
What do people realistically do with these? It's too slow and du... not SOTA-level for coding. It's way too slow for video. I tried simulating an "Fable herding Qwen subagents" and it takes much longer and delivers a much worse result than Fable/Astra alone.
-
Yeah that can't be right, my M5 Max gets almost those speeds and certainly lot faster than what they're claiming the M3 ultra gets. Maybe they didn't have the model setup right or were running it with some unnecessarily high quant (>=8bit).
-
On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.That’s like 12 years worth of OpenAI Pro subscriptions
-
That's a dense model. Of course it will do worse.Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
-
A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
-
The model being tested is 18k as configured.I didn't expect this to make the 5090 to look like a good deal.
All 6 developments of M5 Ultra Mac Studio reviews call it the best machine yet… →
Hacker NewsGoogle NewsMastodonNewswiresXBluesky