Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
77% Positive
Analyzed from 12705 words in the discussion.
Trending Topics
#apple#more#ram#mac#local#models#run#years#pro#memory
Discussion Sentiment
Analyzed from 12705 words in the discussion.
Trending Topics
Discussion (510 Comments)Read Original on HackerNews
The issue was the abysmal marketing and "me too" attitude Microsoft had (and still has).
Apple never came first ... but often just at the right moment and had marketing skills to make it a new trend.
Sure, MS launched PDAs and phone-ish devices long ago, running Windows CE and whatnot but they were awful. It's absolute bollocks that the iPhone was a splashing success because of Apple marketing. It was a splashing success because it worked spectacularly well. Random non-tech people would randomly pull out their newest purchase to show their friends. "And now it's a notepad!" "Look and now suddenly it's a calculator!" Sure, your awful HP Tablet had all that, and a call function, well before. But it sucked. It felt like using a computer while squinting, and not like a magic calculator that can turn into a notepad and then into a phone and then into an iPod.
The iPhone was a success because it worked so well. And it worked so well because the technology was there - in part because they invented it and in part because they had the taste to not bring out a shit product but wait a bit instead.
Newton.
But it had nothing to do with the iPod/iPhone release in his story.
(story starts ~1:15, but I highly recommend the entire talk)
That, or I can’t read.
There are emails unearthed in various lawsuits where you can read Bill Gates screaming at his subordinates: "why the hell can't our partners like Sony and Creative create a similar device? Give them all, give them early access to everything, work with them". In the end MS felt compelled to make their own.
Creative's Muvo^2 already was the poor man's iPod with surprisingly good audio quality as well.
In some ways, anyway. Never owned a Zune myself, but a university classmate did and I was shocked by how poorly it handled non-Latin languages… she had a ton of Japanese and Korean songs loaded onto it, and their metadata all displayed as "missing character" blocks. She used it a lot like one might use an iPod Shuffle despite it having a nice screen because the only way to tell what was playing was by hearing it play.
By contrast my 4th gen B&W iPod which was about 5-6 years older handled unicode just fine.
https://youtu.be/ud6rwVkbovA
And they could listen to it 3 times in 3 days
I assume the M6 will take the crown back and then a few months later Intel/AMD will release a new chip and take that crown back again. That's the state of the world we used to expect, but it's a state that has been missing ever since the release of the M1 in 2020 until Intel finally caught up again this year.
When plugged in... This caveat is so enormous it should almost be legislated. If your computer use is at all portable, a computer that scales down to 20 - 40% of GPU power when unplugged is an enormously significant factor. So far as I'm aware (could be wrong about arm devices?) there's no non-apple laptop that operates at 100% speed on the road.
Hold up, in what metric/benchmark? Feel personally like we life in an age where, no matter the SOC vendor, something high performant and efficient is offered, so seeing a claim that any vendor, be it Intel, AMD, Qualcomm or Apple, is consistently outperforming another, I'd like to get more context on that.
Intel were x86, Apple Silicon is ARM-based. ARM-based chips are more power-efficient.
Also, it's built on a much smaller process. 3nm, if I'm not mistaken, older Intel was something bigger than 10nm. Heck, if you take a really old Intel Mac, you have something like 65nm process, which is much less efficient than 3nm.
Here's a random benchmark I found on the internet (literally the first thing on Google, you can find more if you want) https://www.cpu-monkey.com/en/compare_cpu-intel_core_i7_1065...
Of course you can beat the entry level MacBook Neo by comparing it to larger, more powerful, more expensive laptops.
The M5 is an entire family with a range of performance. Intel/AMD have done a lot to improve performance but they’re not beating the high end M5 chips on performance or battery life yet.
It's also not hard to find a Windows laptop that beats the M5 on battery life. Single-thread performance is the only remaining measure where the M5 is king.
The real question is whether all these Intel Windows laptops spin their fans at full speed when doing absolutely nothing. That’s something I can never go back to. I do some light gaming on my M2 Pro MBP and it gets hot when trying to push 120 fps. But my Lenovo Legion (that’s now collecting the dust) is so loud I could hear it through headphones.
Intel had to reduce the frequency of their processors when running AVX2 instructions and the AVX2 frequency of the processors were non-disclosable to anyone.
Also, benchmarking Intel processors and publishing these numbers were forbidden in some cases. I don't know whether this ban is still in effect.
x86 processors can't keep up with the ARM processors TDP and thermal profile wise. So they slow down a ton when running on battery. See Jeff Geerling's last video on Apple Neo vs. some Intel laptop. It's as "efficient", but slow as a newborn tortoise learning to walk when unplugged and trying to get the most endurance out of the battery.
My M1 Mac gets almost 2 days of low-intensity use after ~6 years of use, and it got warm once or twice because something ran away in the background for tens of minutes.
But talking this authoritatively on something without doing the reading, that's grating: https://chipsandcheese.com/p/arm-or-x86-isa-doesnt-matter
- Apple: We have matched Intel in CPU performance thanks to this new PowerPC! (Shows ad that uses carefully handpicked benchmarks to suggest that the PowerPC is actually faster when it really isn’t on average)
- Intel: Oh, we just found a 30% clock rate increase in the pocket of our other fab pants.
- AMD: Hold my beer, I have the DEC Alpha guys making an x86 CPU… How about 64-bit while at it.
5.25GHz Pentium 4s, intentionally lower binned Athlons, the era your CPU got obsoleted the moment you booted it for the first time.
I don't remember early 90s much. I was too young back then. I don't remember much stuff from that era. But late 90s, early 2000s.
Oh, boy.
P.S.: AMD64 was a great sucker punch though. One of the professors in our university rejected to believe and got mad when he learnt that Intel licensed AMD64 from AMD, heh.
But then you started to see the cracks. New competitors would launch new products with very slightly better metrics than Intel's older stuff, just to be, heh, meep-meeped at the next press conference. But the overlap was real, if small. And it grew over time until everyone looked up around the 5nm node and realized Intel had lost.
That's where we are right now with Apple. "Funny in a way", sure. But history says this is more likely to be the beginning of the end. Everything goes in cycles.
Now we boot an embedded microcontroller (or CPU) which boots the main CPU which boots another OS semi-persistently to boot the main OS (if it's allowed).
Sometimes there are other processors needs to be up to allow processor to continue booting as well (these are mostly servers, but eh).
I'm very very surprised that Xiaomi matches Apples speed even with the newest release, its not diminishing Xiaomis success.
This in-turn later on, AIs will train to make pro-China comments as AIs train on these.
They got the sheer man-power, and with AIs it's even easier.
Apple is the second richest company on the world (which doesn't need help/protection?!) and they have experts in chip design.
Xiamoi is some random chinese company not known for high end chips and was able to catch up impressivly in a short period of time.
This fact doesn't get funny or wahtever just because apple brought out a new chip today.
I'm not a fanboy for any of it and do not care.
I enjoy it because of progress, not because of Apple.
I yield the floor to no one when it comes to pessimism, but that's incredible.
Fab capacity is being bought online; there’s just lead time.
Noticeably greater intelligence is being achieved at the same number of parameters (see: Qwen3.8).
I think the future will be bright, it might be a matter of time. And for tinkers, a used Epyc + DDR4 server can be great fun and epic value.
However a problem exists with scalper bots. They're going to fight tooth and nail to keep prices artificially high even when supply increases. Same with the hyperscalers blowing "free money" at eating supply as a strategic position against smaller competitors.
Basically everyone that makes memory is building new fabs, meanwhile I'm not sure how much longer AI datacenter demand for ram will last. I think the decrease in AI ram demand and the new fabs will likely coincide leading to a collapse in pricing.
That is, of course, assuming the memory manufacturers don't pull their favorite trick and collude.
Could you provide more details about the Epyc + DDR4 server?
The above is a standard project management problem. We do this for lots of industry all the time. There is every reason to think you can get a new factory running in 5 years.
Note that I said 1 factory above. Some of the special machines we don't have the ability to make them fast enough to do 2 (I don't know the real number!) new factories in 5 years. Existing factories are using most of the special machine capacity to replace machines that wore out on the way - this can be corrected as well, but it adds another year and the expenses are much larger. Realistically though 1 new factory is likely enough.
https://transportgeography.org/contents/chapter3/transportat...
...
> I don’t think anyone truly knows when, but it will happen.
Do you know what cyclical means ... ?
You'll just have to be careful to match the enclosure to the ports on the system. The base-model M6 Mac mini still uses Thunderbolt 4, so a USB4v2 enclosure would be wasted.
Compared to the prices of storage today, when people are presumably buying the product, it's actually a bad price.
Ok I'll be that guy. It's pretty easy to figure out if you're talking to an LLM now we know it's tics, failure modes, jailbreak techniques etc
E.g how is the perf/$ vs Wildcat lake
Is there anything better now though?
All I see from AI, is an amplification of the enshittification of the internet.
And people being even more alone.
- Extreme poverty has dropped from 30% to under 10% globally. - Child mortality rates have dropped in half - Internet access has exploded from 10% to 70% - Solar energy costs have dropped 90% - Cancer death rates have declined by 30%
All of these massive improvements in less than 30 years.
While there certainly are issues to solve, and if you simply follow journalism you may think the world is worse off, but for many, their lives have been significantly improved.
It's a Substack that reports good things happening around the world, divided into sections like "Conservation and Restoration," "Climate and Energy," "Medicine," etc. And they also give part of their profits directly to projects in those categories.
(I'm not affiliated with them, I'm just a subscriber.)
But my personal observations of AI is that it's producing more and more of the same stuff and not moving the front forward much. The human innovation and invention seems to be lost.
And the median American is also richer in real terms, both in terms of wealth and in terms of income.
Of course, that all assumes that you use a reasonable measure of inflation that’s stable, well-designed, and applied methodically and consistently over many decades.
Alternatively, you can cherry pick data points and go based on vibes, which lets claim whatever you want!
That's one big plus.
These times are exciting and rough seas make good sailors. Find your path forward.
I'd rather go to the library and read a book.
"Find your path forward!" he shouted with glee, as he ran toward the cliff.
https://en.wikipedia.org/wiki/ELIZA_effect
It turns out the limiting factor isn't how sophisticated algorithms are, it's how gullible humans are.
No AI would pass this test with experienced judges.
Though we're pretty good at sizing up a person's emotional balance/maturity and competence at familiar tasks. So maybe have an old blacksmith watch the AI/robot interact with horse owners for a while, then shoe their horses, and see how well it does.
https://commission.europa.eu/news-and-media/news/safer-and-m...
In contrast, humans tend to paste me the same barely-relevant macro over and over, no matter how much time I spend explaining my issue.
https://www.psychologytoday.com/ca/blog/the-digital-self/202...
I kinda have to link it now, so uhh here's a random PDF: https://www.hec.edu/sites/default/files/documents/Computing%...
And the "popularized" version is faulty also since it uses an ideal, abstract human judge (like the "spheroidal economic agent").
But if you want to add declinations to the said popularized image of the Turing test, you may add Maxim Lott's IQ tests at trackingai.org . Between the end of 2024 and the beginning of 2025 LLMs reached an equivalent IQ of 100, for example.
ELIZA beat the Turing test and then everyone forgot about it. Humans are just really terrible at recognising robots.
I think there are elements showing lowering of performance and expectation.
https://en.wikipedia.org/wiki/Turing_test
I've tried [1] and I almost 100% detect which is the AI. I really want to convince myself I have failed, does anyone know of a better site/resource for this?
I know it might be moving goalposts but I would consider AI to have passed in a well and truly undisputed manner when [2] is resolved.
But in a more practical sense, if AI can impersonate humans so well today then why are state of the art frontier models so obviously AI when they create PRs, commit messages, documentation, etc. Are the companies deliberately making them unnatural?
[1] https://turingtest.live/
[2] https://www.metaculus.com/questions/11861/date-when-ai-passe...
So, on the mini the RAM upgrade runs at 25$ per GB on all tiers, the same as the Studio therefore the upgrade to 512 will probably cost 6400$.
The fully maxed out Apple Studio then will be 24699$. It's 17199$ if you don't upgrade the storage(1TB).
Nevertheless I itch to have one :)
EDIT: or buy AAPL. If I had bought Apple stock instead of buying a Mac LC II in 1992, then I would have about $2 million in Apple stock.
Depends on where you live. Also keep in mind population decline so long term, im not so sure. Prices already dropping in less desirable places (everywhere outside of most blue areas) (Assuming you are talking about the US sorry if not)
For general inference there’s no ROI that makes this work vs subscriptions.
25k for computer now, plus 9-10% sales tax, plus operating cost, plus time and cost for R&D tinkering with models, harnesses, and infra (assuming highly capable engineering talent that can get paid for your human inference) vs a HEAVILY subsidized subscription at 200 per month with free R&D has a pretty long ROI (15 years?)
At API costs, it’s like 6 months if you’re heavy on inference. For training, specialized models will have their own ROI that makes this worthwhile. Then debate renting capacity and the platform to choose
Unless you need privacy for your inference this instant, paying for credits can get 80 to 90 percent of people everything they need.
Of course if you do need that privacy, then forking the $25K over to Apple is a no brainer.
… I wish I hadn’t just calculated that.
In RTP (NC), a ~$400k house at 5% down is $20k
I wouldn’t recommend buying any bare metal unless money is a second thought or you can fully deduct the price.
Most often in the end you pay half the price then. Depending on the write offs you could even make some bucks out of it.
Or buy and lease. Under certain circumstances the hardware costs you nothing.
But you need money to save money. And a company.
Put together a similar build with a couple of rtx 6000 Ada cards and Apple's price tag suddenly looks pretty damn reasonable
They’re very different things.
The more logical argument to me is that Apple uses its upgrade price points as more than just direct BOM and rather as a proxy for things that are amortized across all their sales like support/warranty/etc so higher SKUs subsidize the costs of the lower ones.
The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig.
Hybrid ‘small local’ and buying inference is the way I choose.
A 4 GPU linux box with 3090s, which are $1500 a piece right now, will blow this thing out of the water. Even 2x3090 rig will run most of the good local models like Gemma4:31b at 100+ tok/sec
The VRAM of the GPUs are MUCH faster than the unified ram within Apple Silicon. The only difference is the initial model load, which takes longer from disk to VRAM due to PCIE limitations, but once the model is loaded, GPUs can prefill and and generate tokens way faster than any Apple Silicon.
So is your desire to upgrade to mac because you just aren't aware of how to set up a GPU rig, or is it something else?
"According to reports from Bloomberg, Apple will be skipping its M6 Pro, M6 Max, and M6 Ultra chips to accelerate development of the M7 chip. That means the only chip to be released from the M6 family will be the base M6.
The reason for this break with tradition: AI. Apple had been planning major neural-processing upgrades for the M7 family and ultimately decided those improvements were important enough to justify accelerating the next generation rather than completing the M6 lineup." https://9to5mac.com/2026/08/08/apple-m7-chip-heres-why-it-ma...
I'd skip M5 and M6 chips for LLM work and wait for a year for M7.
I believe that the CPUs are actually limited by ram bandwidth more than the neural engine right when it comes to LLM processing?
Maybe the M7 introduces something new to get around the current ram bandwidth problems on the non-Ultra chips.
Please say more? Is it because it is a one-time cost, unlike a recurring subscription of Claude/Codex?
I wasn't able to debug network errors (restartin my Mac worked), Metal was missing low level disassembly / debugging tools (there is some hard to use UI), but the worst thing was the inflexible windowing system.
Even getting all the window handles on all screens/desktops with their titles and programs is impossible.
I just decided that I move to Omarchy 4 (basically Hyperland + QuickShell) + NVIDIA GPU, and I already was able to customize it more than my Mac in years.
I will miss Apple's hardware for sure, but not MacOS and the missing hardware documentation
It’s a bit depressing because it means that if I ever feel forced to switch my daily driver, it won’t come without a dump truck load of friction, frustration, and lost productivity, which I’ve validated by using the various Linux desktops on secondary machines.
It's still not well integrated of course as those plugins are from different people, but I at least don't feel powerless as I know I can make any change easily.
For example when using PyTorch I wanted to try to speed up my NN kernel by 2x by just using half precision and haven't noticed any speedup at all. Also I was missing the easy to use GNU tools that had to be mixed with Apple's tools.
I loved using Arc browser as well, and I'm missing it, but I guess I will do without it somehow (Chrome's vertical tabs are just not the same).
My main program missing from going back to Linux was ChatGPT Desktop, but now it's there.
I just checked out Hammerspoon, I'm happy for you that you wrote it, and looks great, but it has the same problem that I had: for security reasons Apple stopped allowing the window APIs to get all important information on other workspaces. You can only do it with Accessibility API. I was trying to fight with it but have up.
[0]: https://zen-browser.app/
I ordered an ASUS Zephyrus G16 with 5090 NVIDIA card + 1.9kg (quite an overkill, and I know that I will have to limit power output), but hasn't arrived yet.
But what's fun is that I love QML+QuickShell with its hot reloading, Hyprland with its Lua support.
With AI nowdays it's just so easy to do deep UI changes that wasn't possible a year ago.
Also, a lot of companies are looking at how to run capable models locally to cut some of their (massive) cloud AI bills. An easy answer is worth a lot to them.
What makes this expensive & sell well is it's not very fungible at the moment. Where else are you going to get 512 GB of high speed memory with a well supported accelerator attached that you can throw in the corner of anyone's home and not really have them notice? There are plenty of lesser options, plenty of noiser/power hungry options, plenty of harder to support options, but not really something in direct competition at the moment. Even the next rounds of the integrated AMD/Nvidia solutions are only targeting 196 GB of much slower memory and compute.
Closest competition I see right now are stacks of 2-4 connected DGX Sparks, similar lowish speed high mem, and about the same cost/gig.
For me personally, not quite that valuable yet, but I think it's getting there quickly. Deepseek V4 Flash massively increased the value of local AI to me, to the point where it's displaced most of my Claude Code usage, its upcoming vision enabled version should bump it further, and it's only going to get better from there.
It's a lot faster, but a lot of it is also feeling free to discuss things I wouldn't be comfortable sending to Claude, with the idea that that info is now theirs in perpetuity. I got my genome fully sequenced recently (it's cheap now!), and I get a battery of blood tests every year. Wouldn't do processing on any of that with Claude, but local AI? Totally great.
And if I was running a company with a large cloud AI bill, I'd probably buy a wheelbarrow full of these macs. Cheaper, but also a more solid/predictable base to build on.
Put another way: If $25k is the full extent of the start up capital costs, and operating costs are very low, that is a much cheaper business to start than most! The question is whether this is actually a useful model for a revenue generating business. I think that remains to be seen.
My guess is that there will be a few hits (which we'll hear a lot about - especially when someone actually pulls off "the first single-person unicorn", which I do suspect will happen someday) and a huuuge number of misses, which we won't hear much about.
AI based tools are very useful here - thinks like object removable or cleanup etc, not just AI generation.
For example Apple mentioned performance increases for https://learn.foundry.com/nuke/content/reference_guide/air_n...
This is, sadly, probably a foreign concept to a lot of people who have only worked at companies where hardware purchases are viewed as something to minimize and everyone is stuck with the same low spec laptops that the finance department picked out. At companies where someone might have a legitimate use for a $20K machine, their fully loaded costs (not their salary) are $300K or more, and other teams like sales are spending thousands of dollars per week on things like travel and hotels for their job, spending $20K on a computer that’s going to last several years is not a hard choice.
Even using multiple windows in parallel for as many as 5-10 hours per day, I find that I am not fully using my claude max (20x) and chatgpt pro (20x) accounts. I can for sure use up the claude max account, but chatgpt either gives me a free reset before I run out of tokens or I just fail to use the full quota. The quota for Sol seems like 10x that of Claude Opus at the same level, and forget Fable, you can use a 5 hour quota in 20 minutes.
But lets do the math:
Lets say a 20k workstation can run 1 inference at a time at the same speed you get with Sol hosted by openai (big assumption) and run an equally capable model (big assumption).
Each month this gives you about 100-170 inference hours on a Sol 20x Pro account, and 720 hours (if you utilize 24/7) on the workstation.
Assuming a 36 month amortization before the workstation has to be replaced due to no longer being able to run frontier models or is too inefficient due to electrical costs or what have you:
The monthly workstation cost is about $550 capex and $150 electricity -> $700/month
You would need about 6 Pro accounts to reach that capacity, which would cost you $1200 a month.
But this fails because:
- You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day.
- During work hours you are capable of utilizing more than 1 concurrent session. 6 Sol accounts would support as many as 20-30 during working hours, not all the time but if you could burst to that many (don't forget sub-agents and agent directed parallel agent workloads).
- In 1 year the cost of Sol level models is likely to cost a fraction of what it does now.
this leads to:
I have agents running 24/7 doing research, in fact I would argue this how they will be used for most programming tasks in the near future. For chatting, I agree local inference makes no sense. But for tasks that run continually, I'm not so sure. Personal computers took a while, local inference will too, but I think it will happen.
isn't the whole point of all this ..... agents? isn't that what literally everyone is always clammering about in these threads? in which case the workstation is useful 720 hours out of 720 hours.
If Apple didn't sold these things they wouldn't make them but, also the level of marketing that Apple is talking about for AI is basically the new group they need to capture because the ones I just listed are already buying Macs and or easily to motivate with the other obvious CPU / GPU performance upgrades for code compilation, faster memory and video transcoding.
to answer your question : looking at the aftermarket availability of Apple's prior best and brightest : practically no one buys them.
"people here buy them" , well, 'here' is one of the most affluent groups of people in the world.
They're available as movie and television set pieces (undoubtedly disappearing into the home of someone close to the staff post-production), and for administrative/boss types that can slip the cost into a ledger somewhere that few will ever see.
It has been a hobby of mine every few years to check out the apple site and see how big I can option a machine. My record was when I was in high school years ago and was able to option some pro studio-ish apple desktop thing to like 61,000 usd out the door.
For one thing, you can’t tell from a movie what the specs are. A $999 Mac Studio looks exactly the same as a $20,000 one.
For another, Apple updates the industrial design on their products so rarely, a 6-year-old Mac, iMac or MacBook also looks nearly indistinguishable from a brand-new one.
Its cheaper than Nvidia AI hardware.
That may not be many people, but there certainly will be some people who want to do that, and are willing to pay big bucks to do so.
It’s when self hosting and local hosting was the norm, and why it’s also starting to come back.
There will be workloads that can never touch a public cloud, and for it solutions like this are an option.
I've often felt there is tremendous value locked up in underutilized old computers. It would be interesting to see Apple in 3 years offering compute as a service using lease returns (or more likely, partnering with someone else to operate it (perhaps exclusively in secondary markets like China or India, to address political demands for local siting or jobs). Apple is in the best position to work around or even gap-fix older software/hardware limitations in a controlled environment, and now they can do so without cannibalizing new hardware sales.
Apple computers tend to have excellent resale value, and Mac Minis/Studios have the least depreciation of them all. I understand the benefits to both taxes and cash flow, but boy is Apple winning big on those lease offers for Studios.
consumer electronics has literally never been an asset.
> Apple computers tend to have excellent resale value
do you think the new leasing category might change that? hmmmmmmmmmmmmmm
edit:
https://www.reddit.com/r/LocalLLaMA/comments/1vxzg6v/apple_i...
For which configuration, though?
not getting on that bandwagon but wasn't that not the most demanding game as its a just a nonstop cutscene.
Mac Mini + MacBook Neo w/ ssh can be a better setup than MacBook Pro for many people.
Nevermind that Apple still insists on providing base systems with only 512GB of non-upgradable SSD. The equivalent spec mini (4TB/64/10G) is almost $5k for half the VRAM. Not as good a CPU compared to the M5/6 but you also get 20 cores and full CUDA.
Why choose the worst game of the year? Just because it's made by the daughter of Larry Ellison?
I’m not claiming this title was made by GenAI. I’m saying something much more insulting: that the human artists who made it have no talent.
Which one is it you can run local models on? I suppose the NPU only.
In US its $4000 upgade so $25 for 1GB.
Also:
> 512GB memory option for M5 Ultra coming late October
If you want to comfortably afford this gen you had to trade options on memory stocks...
Got the recommendation from these articles: https://www.xda-developers.com/qwen-3-8-27b-reverse-engineer... https://www.xda-developers.com/lenovo-thinkstation-pgx-revie...
But haven't had a chance to try it myself.
For tinkering and learning, it's been great. Tie it into something like Hermes and you have a pretty powerful AI assistant in a box. And when you need to step up your model, you just do something like OpenRouter and it makes it pretty easy.
The problem I'm having now is that no models are targeting RAM of that size. Everything is either much smaller, targeting laptops, or much larger, targeting hardware well out of reach of enthusiasts.
Please, AI people, start making models targeting 128GB machines again. The last interesting one was Qwen 3.5 122B.
Something like this would give you three concurrent sessions, each with 240k token context:
Should bench better than Opus 4.7.
You might expect the M5 Ultra to produce 50 t/s from Qwen 3.8 27B with a good context length.
Interesting times, to say the least!
I plan on maximizing my residual student benefits, and taking advantage of education pricing.
I edit 4K ProRes and H.265 footage, sometimes with multicam (up to 4 streams) and color adjustments, titles, etc. It's only after stacking 3-5 effects before things can stutter, really.
Or if you try doing something CPU-intense in the background _while_ running some heavy creative software. I just don't do that.
That being said: you might want to consider at least 2TB storage. Having 256GB RAM to fit models is less useful when you can only store a 3 or 4 large models on your disk. The M4 -> M5 transition doubled the drive bandwidth (at least for the MacBook Pro, I'd assume the Studio is no different) making them extra nice for loading large models. Then again, you can always add an external NVMe drive over thunderbolt, so the tradeoff is less straightforward than with a portable MBP.
The prices are ridiculous though. I may just keep rolling with my Windows 10 setup.
Today I entered this new M4 mini's serial as a trade in for the M6 base mini, and the retail price went from $1249 CAD to $824 "once trade-in received". That's the automatic estimate, so basically I'd be paying twice for an upgrade...
That's wild!
This may depend on the size of your library. I tried installing Jellyfin on a Synology NAS, which runs Plex just fine, and it ran so poorly it was basically unusable. It “worked”, but it was painful.
Edit: never mind, 2.5G is now the default, so you'll probably need more than 40 people to saturate that with streaming video.
Just amazing engineering push, the competition got the message and we benefit.
I can't believe that Apple still comes with this bullshit like 32 GBs is a lot. It's a lot for video memory - vRAM, but not RAM.
Also: "a staggering 1.2TB/s of unified memory bandwidth" -- yay, the GPU has reached the year 2020! (I'm a bit bitter that my M4 Max is near useless for local LLMs because of its low memory bandwidth.)
A m4 mac mini is better than al of these per dollar, msrp adjusted.
Hopefully by the end of the decade China figures out manufacturing at scale and fixes this.
Apple never needed to participate in the AI race to zero. Because they were already at the finish line years ago building their own chips that can run large >100B parameter AI models locally.
It's possible that they're working on their own LLM that's going to work very well on their chips, and possibly outperform anything out there when they do release it.
10 years ago 32GB ram laptops sounded too much. 8 was enough. These days even I would get that much ram since it’s soldered. 64GB is higher end.
In a few years we should see such high end hardware commonplace. Working with a local LLM to get work done is the ideal way to go which has mostly hardware limitation as of now that gets solved in due time.
Ten years ago I got 64gb of ram in my laptop, same as I have now. I bought both for business and personal use. System ram capacity hasn't changed much in 10 years.
It makes me curious how old you were 10 years ago.
The iPhone 15 was almost entirely marketed based upon AI (I would say fraudulently so, advertising features they still haven't delivered), and a huge portion of the OS work was on local AI or AI integration.
And for that matter Apple has been dumping enormous sums into their own AI development. Their failure to have a lot to show for it doesn't void the fact that they tried really, really hard.
It's bizarre how often this "Apple sat on the sidelines and let the AI people fight...so smart!" narrative appears on HN. Apple hasn't gone down the path of spending hundreds of billions on nvidia GPU data centres, but they absolutely tried really hard to matter in AI.
Yep, Siri AI; they’re doing it in public.
But you get a generic computer and much more RAM.
And you lose a couple of organs.
It would take Apple one or two engineers to make Linux life much easier on macs. But Linux is outside their walled garden so it's ignored.
My Mac Mini is strictly a headless server for llama.cpp.
I use a Linux workstation.
If I were limited to use Mac hardware , I would install Linux in VMware Fusion and work from there.
My old 3090 is typically significantly faster (almost 2x token/s) than my M4 Max 128GB machine, as long as the model fits in the 24GB of VRAM.
In most situations it's a better idea to just buy tokens. But there are definitely cases when that's not an option. And then a machine like the M5 Ultra can allow you to do things locally for a fairly limited budget. And in a simpler package to manage than a machine with multiple GPUs.
I mean this is not nvidia based right? It's all custom? So we can use it under Asahi perhaps?
I want to get something for my company to run local models, wondering what would be a good option.
Linux runs very well in a VM on macOS. There are many good options for this, some free and open source (QEMU, UTM, Lima, Colima), some proprietary (VMware Fusion, Parallels).
But Linux in a VM doesn't get access to the real GPU, so model performance is limited. Those running on the CPU perform well, and those needing the GPU don't.
However, macOS on M-series macs is excellent for local models. (Maybe not as excellent as a box full of the best nVidia GPUs, but still excellent).
So if you're getting Apple hardware, like Linux, and want to run all of it locally, a fine setup for a machine to run local models, with agentic characteristics:
- macOS running one of the many local model runners. I used to use Ollama and Whisper, and now use llama.cpp instead of Ollama. Others use LM Studio, oMLX, etc. Provide HTTP endpoints to access the models.
- Linux in a VM for overall control and orchestration, with standard VM settings, and bridged networking so it appears as its own machine on your network. Also, in here provide a robust shared file server for shared state. Use this VM as your desktop and primary access to the machine, if you like Linux.
- Linux in a VM to launch ephemeral, volatile containers, with the containers using a memory-only tmpfs overlay on top of a read-only Linux filesystem in a VM disk image, with tools in this filesystem. Alternatively, a writable Linux filesystem in a VM disk image, with disk buffering set to use macOS host buffering and discard fsync requests. These settings optimise for container disk performance for data that's only ephemeral which will be deleted soon or on system shutdown. (You can combined both VMs, but need to use two VM disks to get equivalent behaviour, and be careful about VM disk configuration of the two disks.)
- Containers spawned within that second Linux VM can be spawned very quickly and run quickly, so are ideal for LLM agents that need a quick sandbox. These sandboxes generally run faster than a macOS sandbox, despite being on the same machine with VM overhead, because Linux is faster at some things. Teach the LLMs to store files and memories they want to keep in the shared file server.
I just want my butt ugly repairable beast machine to do the same trick. Why is my ram not unified? I have an iGPU in my server, but it can't access the 64 GB ram (I got last year for 150 euro) directly or something? It's on the CPU right? Why did only Apple go for this architecture? So many questions...
Can we please kill the xcode. It is worst pile of garbage I have to use just to develop ios app.
Is this a joke?
>Additionally, M5 Ultra features a massive amount of high-bandwidth unified memory, up to 512GB
Now we're talking. But at what cost?
"M6 also introduces a Dual 16-core Neural Engine, providing up to 2x the peak compute over previous generations to make on-device AI workflows run even faster" .
256 memory gets you to like 11k. So like 15-20k.
and 512GB is so 2025.
Give us 1TB version. Where is the competitive spirit?
To be fair, where is the competitor at this form factor?
It's the Hardware, Software and Services in combination. None would work without the other (to reach the scale apple is)
More specifically, they're a hardware dongle company first
FWIW, Ollama, LM Studio and Lemonade (and oMLX) also wrap Apple's MLX framework.