Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

91% Positive

Analyzed from 983 words in the discussion.

Trending Topics

#mac#ultra#models#local#hardware#studio#apple#run#numbers#more

Discussion (68 Comments)Read Original on HackerNews

simonw•about 2 hours ago
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

  Qwen3.8 27B tokens/sec generation speed

  Prompt size    8K    64K   128K   256K
  RTX 5090 PC    59    51    44     n/a
  M5 Ultra       48    39    32     24
  M3 Ultra       31    23.5  20     15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
RationPhantoms•about 2 hours ago
Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090.

Maybe Apple is an acquisition away from changing that balance.

gpugreg•about 2 hours ago
Those RTX 5090 numbers are bad. Should be over three times higher (and should use NVFP4).
peri-cl•about 2 hours ago
Those are some incredible graphs (at the anchor you linked at). That's a gigantic improvement in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3, locally.

/meta Here's a (uBO style) CSS filter to stop those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
redox99•about 2 hours ago
A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
srcreigh•about 2 hours ago
This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

crossroadsguy•38 minutes ago
My mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That's a shame. If only I could install an alternative OS that uses very little amount of RAM :-)
tempoponet•about 2 hours ago
While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.

This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.

sajithdilshan•about 2 hours ago
On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.

That’s like 12 years worth of OpenAI Pro subscriptions

112233•about 2 hours ago
Hard to guess, it can go either way. If you will need to be in a syndicate to use non-sterilized models, that mac makes sense. But if there is mandatory registration of personal cyberarms, you risk going to mines once they check you purchases. You could try to play normie and pretend you simply wanted to show off, by keeping your actual work on external disk, but that leaves traces on system. Counting on someone in the Gap renting you gray iron works as long as you can swap credits. Still, this gear is tiny. Put it in your e-car, with uplink, and leave it at uncle's farm. Discreet.
simonw•about 2 hours ago
Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.

Plenty of other reasons to get excited about it local AI, but I don't think cost is one of them.

geodel•about 2 hours ago
Agreed.

Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.

vardump•about 2 hours ago
I hope that was sarcasm.
akozak•25 minutes ago
"a total cost of $0" Uhh ... how much is that hardware?
liuliu•33 minutes ago
When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
ApolloFortyNine•about 2 hours ago
The model being tested is 18k as configured.

I didn't expect this to make the 5090 to look like a good deal.

theplumber•33 minutes ago
At this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing
SamuelAdams•44 minutes ago
I think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.
addaon•about 1 hour ago
Ordered one for OpenFOAM. Excited for it. Will be nice to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.
Advertisement
12kaj2•28 minutes ago
The Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.
prmoustache•7 minutes ago
The year of linux on the Desktop was 26 years ago for me.
BatchJob•about 1 hour ago
If you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.
devy•about 1 hour ago
This dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!
snarfy•about 2 hours ago
$12,299
WarmWash•about 2 hours ago
>Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster?

Ehh, the actual elephant in the room is:

"why bother with local AI at all when you can lease a GPU for $5/hr?"

To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.

sghiassy•about 2 hours ago
Imagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI
whalesalad•about 1 hour ago
For 99.99% of people, spending 15 grand on a Mac Studio just to run Qwen 3.8 locally is a non starter.
jmull•about 1 hour ago
It's not the M5 Ultra itself, but the M7s or M9s that will do the damage.

99% of people will use whatever AI is free. The sophisticated, heavy users that are willing and able to pay a lot of money the ones that will be interested in controlling their inference bills.

Today, the sweet spot where an M5 Ultra makes sense is tiny. But we might expect that to grow a lot.

sghiassy•12 minutes ago
Yes, but in 7 years?
beastman82•25 minutes ago
at 15 tok/s
kokonokko1337•about 2 hours ago
> "It also happens to be a Mac, with an operating system that looks nice and doesn’t suck"

Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.