Back to News
Advertisement
ccarloslfu about 14 hours ago 92 commentsRead Article on github.com

DE version is available. Content is displayed in original English for accuracy.

I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift.

It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next

Advertisement

⚡ Community Insights

Discussion Sentiment

83% Positive

Analyzed from 2000 words in the discussion.

Trending Topics

#more#don#model#project#models#memory#mlx#readme#context#before

Discussion (92 Comments)Read Original on HackerNews

embedding-shapeabout 14 hours ago
> Hugging Face is the bottleneck, not your link.

README could clearly make use of a cleanup, seems to be more like a session log dump now than a good introduction to the project for a new user. Maybe try something like "Remove anything from the README.md that wouldn't be helpful to someone who sees this project with zero context, for the first time. Rewrite all paragraphs and sections to be concise and remove all fluff, leave only important details new users must know before using the project".

trollbridgeabout 10 hours ago
I don’t want to be a cranky codger, but I dont get why 5 minutes of work cleaning up the README can’t be done before posting to HN.
ricardobeatabout 8 hours ago
> Run Qwen3.8-Flash-Next on a Mac that can't hold it

This is the first line of the README. I can't believe people are becoming ok with this, and I'm 100% on the AI train.

Eufratabout 13 hours ago
I hate this AI style writing because since it doesn’t really understand flow, it’s being inserted in irrelevant places and it is extremely irritating to read.
carloslfuabout 13 hours ago
I feel you! fix incomming
Eufratabout 13 hours ago
For what it’s worth, this comment was not targeted at you, but rather the model kinda forcing it. I get the sense that Anthropic did not think much of this, but it seems to have gotten worse with recent models and it really comes off as a kind of nails on the chalkboard writing style.

I have to image whatever style of writing this was trained on is a lot more pleasant to read and I feel bad for whoever writes like this now being associated as bad AI writing.

makiraabout 10 hours ago
I used similar prompts before. Now I simply say to "remove historical cruft" and results are good enough. It's the model itself that first used this wording, I found it concise.
carloslfuabout 13 hours ago
thanks! I'll do!
mulemisterXabout 11 hours ago
I have a 48GB M5. I don't need to run larger models. I want more context. I've managed to set the context window at 71,680 using Qwen3.8-27B-oQ4e-fp16-mtp. But I want more. Is anybody, with similar specs, able to set their context window higher?
hadlockabout 8 hours ago
We are running 35b-A3b with 264k context (the model's default max) using vllm and the "frog" jinja templates: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates and had good luck. We are mostly running agentic workloads though, rather than coding. 27b has a slightly higher agentic job completion rate (95% vs 92%) but the 3% trade off is worth it because the A3B is sooooo much faster, and we reprocess the other jobs with a different model. Don't sleep on the froggeric templates.

Qwen: Looking at you for a new ~35B MoE! Please and thank you

kamranjonabout 7 hours ago
I am running 3.8 27b at q6 quant with 160k context on a 32gb video card (arc b70 pro) - I quantized the kv cache at q8 - that is the only trick really - works great.
pramabout 9 hours ago
You should try Glimmer MTP. Qwen3.8 27B seems to have weird memory and caching issues on oMLX
ig0r0about 10 hours ago
yes, with qwen3.8-27b-4bit run via rapid-mlx i can get to about 200k
metadatabout 10 hours ago
Nice.. and how many toks/sec?
ig0r0about 1 hour ago
Around 20
prometheus1992about 13 hours ago
It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat
monster_truck5 minutes ago
The laptops definitely can't hang but the minis don't really care. I threw mine down in the basement just to put the heat somewhere else, can tell when the dehumidifer next to it is on because it's a few C lower but that has no impact on performance. I don't think it's ever seen anything north of 70
Baloogaabout 12 hours ago
Now I'm feeling pretty good about getting 10-11 tokens/sec running Qwopus 3.6-35B-A3B Q6_K on an old Mac Pro 2013 (trashcan) with 128GB RAM (DDR3), 12 core Xeon, dual D700s. Arch Linux and llama.cpp.
prometheus1992about 11 hours ago
haha, good for you.
trollbridgeabout 10 hours ago
Anything smaller than a 16” runs into serious thermal problems; even an identically equipped 14” just can’t dissipate enough heat.
carloslfuabout 12 hours ago
interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out!
whartungabout 13 hours ago
I'm hoping to see progress in this space.

Folks talking about how 32G is not enough for local use, but then there's been work like this to empower it.

My hope is that the new 32G M6 will be "useful" locally, possibly because of work like this.

tyreabout 11 hours ago
Yes, but also 12 tok/s versus Claude is so far from comparable. I know that it’s not exactly 1:1, but it’s a long way from an easy trade-off, especially considering hardware prices for high levels of RAM.
trollbridgeabout 10 hours ago
32GB is simply too tight; you need 8 minimum for the OS and you need about 4-8 more for the LLM you’re visiting and kv cache.
carloslfuabout 13 hours ago
yes! I'm bullish on this. there is a lot of work to do. I've been experimenting with pruning, distillation, and retraining too. I'm sure your 32gb m6 will run a badass local model!
jmward01about 6 hours ago
Not a mac/UMA discussion point, but is it time to add additional, installable, DDR5 to GPUs? I can see this as a win/loose. PCIe 5x16 is close to maxing out the bandwidth available from high end dual channel DDR5 now, but not quite. I'm not a hardware person but I suspect putting it on the card could lead to significant performance improvements over using system ram so allowing systems like this, where MOE weights are shed, to get even higher performance than just adding that DDR5 to the system. Bigger models become closer to reality and it provides more of a pathway for developing technologies that take advantage of it. Of course the loose side is that you just put a lot of specialized ram on a card instead of into the system where it could be used for other things. I could see a place for a 16GB card with 64GB(or more) of DDR5 especially if we start seeing MOE and similar technologies really start being designed for this concept.
atif089about 13 hours ago
As someone who is just looking at the theoretical benchmarks of each of these models I'm curious if anyone could share what are the problems (maybe around code) that flash-next was able to solve which 27b was not able to
red_hareabout 11 hours ago
For a local non-coding agent, instruction following and tool use are the most important gains
carloslfuabout 12 hours ago
This is the best I could find: https://huggingface.co/Qwen/Qwen3.8-Flash-Next?utm_source=ch...

About the specifics, I have only anecdotal evidence, but I guess this info can be found somewhere

jacquesmabout 11 hours ago
I love these efforts to get proper models running on lower cost hardware and I think this is where the next real breakthrough will come from. The more efficient this sort of thing can be done the bigger the chance to democratize this tech, 'good enough' is what you need and as long 'top of the line' gives a competitive edge even if it is at a cost there is a substantial risk of the door closing on general computing at some point in the near future. Keep in mind that there is no guarantee that the pendulum has to swing back, it can swing one way and get stuck, and then you're going to have to beg for crumbs from the haves.
cosmic_cheeseabout 10 hours ago
I think there's a very good chance that history will rhyme a bit.

DOS/Windows and PC clones were by no means the best available, but they were cheap, ubiquitous, and versatile compared to alternatives that were either much better at one task but more expensive or better at everything but wildly expensive. They were "good enough" and represented a solid improvement over what many existing computer users had as well as a good entry point for new users. As such they spread like wildfire and became the standard while the expensive alternatives either became hardcore niche or vanished.

jacquesmabout 10 hours ago
SUN Apollo SGI

Though to be fair it was Linux more than Windows that killed them. Dos and Windows were competition for DEC and - ironically - IBM.

ameliusabout 9 hours ago
How usable is 12 tok/s?
pornelabout 8 hours ago
Unpleasant for interactive agentic work. Still useful to leave it to do some work in the background.
c0rruptbytesabout 6 hours ago
so many inference project, omlx already supports all of this and has a 1000 people trying to optimize it constantly
carloslfuabout 5 hours ago
Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX
carloslfuabout 5 hours ago
Interesting! I'll check it out
drcongoabout 14 hours ago
"Disk is the gate that bites first"

AI;DR

thirtygeoabout 12 hours ago
Ha! AI;DR is a great phrase. Had not seen that before
drums8787about 13 hours ago
The never ending gate bites.

How I have come to detest certain phrases.

bogzzabout 13 hours ago
Load bearing gate bites.
drcongoabout 12 hours ago
...the seam.
ErenayDevabout 14 hours ago
how much energy does it consume?
carloslfuabout 13 hours ago
Good one! I haven't measured this. I'll include it!
Advertisement
siris9476about 9 hours ago
32GB dedicated to an N-gram table instead of a draft model is an unusual choice for speculative decoding — what made it win over the more common draft-model approach here?
karmakazeabout 14 hours ago
It seems we could use a new kind of memory that streams the weight data in, like GDDR in reverse.
0x457about 14 hours ago
rzzztabout 9 hours ago
Optane? (Too soon?)
0x457about 9 hours ago
Optane was targeting the latency gap, while HBF targets bandwidth. Given how LLMs work, HBF is perfect for offloading.
carloslfuabout 13 hours ago
interesting!
carloslfuabout 13 hours ago
yes! I guess future hardware designs will have something like that!
kethinovabout 11 hours ago
Next help us normies run GLM 5.3 on our potato computers. Wouldn't that be nice!
kzrdude18 minutes ago
Colibri did that first for GLM-5.2 https://github.com/JustVugg/colibri
jonplackettabout 13 hours ago
Is this going to destroy my SSD?
egorfineabout 13 hours ago
no it's reading, not writing
ElectricalUnionabout 7 hours ago
Using macos on low memory regimes will make it use disk-based swap.

For example, a Macbook Neo (so in theory, something with around 4GiB of free RAM lying around) might eat around 900GB of writes a day while not doing much at all, because it's basically on low on RAM and swapping all the time.

Gigachadabout 9 hours ago
A particularly worrying situation considering a dead SSD will render your macbook usable for parts only.
cromkaabout 13 hours ago
By reading it?
mrobabout 10 hours ago
Reading causes insignificant wear ("read disturb") that likely isn't a problem, but I don't think it's possible to issue pure reads to modern SSDs. The NVMe spec mandates tracking the amount of data read, and this has to be written to the drive. I'd hope the firmware buffers this and writes it at low frequency, but on the other hand, I doubt the firmware was tested in extreme random-read regimes. Unexpected failures from excessive statistics recording could be possible.
carloslfuabout 13 hours ago
I don't know actually. I'll check haha. My best guess is it isn't.
carloslfuabout 13 hours ago
I hope not! this is a new macbook lol!
nikanjabout 11 hours ago
I swear the models are named by the beatbox aliens from the post office in MiB
AmazingTurtleabout 14 hours ago
There are already a handful of repos doing essentially exactly this: `mlx-moe-offload`, `streamlx`, `mlx-moe`, `mlx-flash`, and `deepseek-v4-flash-mlx` - i.e. keep the resident parts of an MoE in unified memory and page/stream routed experts from SSD on Apple Silicon.

At this point I'd much rather see people collaborate on one of these implementations, benchmark against them, or upstream the useful bits into MLX/MLX-LM instead of producing yet another near-identical repo.

The local-LLM ecosystem really does not need every implementation idea rediscovered five times and wrapped in a new README. AI-assisted coding makes producing a new repo cheap; maintaining, benchmarking, and integrating one is the actually valuable part.

carloslfuabout 14 hours ago
I see your point. As an oss defender myself, I agree, however, the spirit of this is to see how fast I can make it. I'm sharing this with the community, which I think is aligned with the original oss spirit.

It's an experiment for myself but I am committing to maintain it. I've been an oss person for a loooong time, way before AI was a thing. Think about it as a new, from-scratch take at it, not as a re-reproduction.

xlaynabout 13 hours ago
Hey carloslfu, kudos from the other side of the internet, don't get down on people nitpicking everything here, experimenting and discovering is part of learning so keep going!, remember this is the place that said dropbox was dumb and could be replaced by a script.
mannyvabout 10 hours ago
I think multiple people working on the same thing is great.

Everyone comes at it from a different point of view, and some approaches work, some don't. And when people do this themselves they learn. Existing projects have their mistakes worked out already.

Maybe one of these people is going to come up with the thing that nobody else thought of because of their experience working the problem from scratch. You may not get that from someone working from an existing project, because existing projects have their approach "baked in."

What all these projects are showing so far is that it's possible to stream from disk, but that the performance isn't ideal. But I'm sure you could take this approach with smaller models and get better performance.

In addition, it's a given that when you work with large data sets performance means organizing the data to take advantage of caches, both disk and cpu. It's not clear how that would work, exactly, given that each run is a not-quite-random walk through the data. The Big Data way is to prebuild all of that as much as possible, which is probably impossible with a big model. But what about a smaller model?

carloslfuabout 5 hours ago
true
brailsafeabout 12 hours ago
This is one of the aspects of this year that I've been finding very grating and wasteful. Collaboration still happens among people with the ability to do so and the technical skills, but everyone else is taking their own helicopter to the top of the mountain, "putting it out there", and there's just a ton of redundant projects that do the same thing.
brcmthrowawayabout 11 hours ago
It's horrible. Every 20-something working on a load-bearing inference engine on GitHub.
carloslfuabout 5 hours ago
for the record, I'm almost 35
Barbingabout 14 hours ago
Vouched especially since OP might have a perspective on this. And readers may want to look up those other repos and compare for themselves.
carloslfuabout 14 hours ago
Thanks for the feedback! I'll create a section with a benchmark and comparisons. This will hold the project accountable and speed things up imo
Barbingabout 8 hours ago
Absolutely! Nice.

(The comments under the parent indicate it was improperly flagged/made dead (maybe could happen just from downvoting?) so glad I hit the Vouch.)

genxyabout 14 hours ago
Why should they do that? For you? You could merge those projects and see if they get traction.
kzrdudeabout 14 hours ago
And there are `Mference` and `SwiftLM` too, I think they are doing the same use case.
dofmabout 14 hours ago
AI NIH
carloslfuabout 14 hours ago
Sorry, I don't get "NIH". what's that?
noir_lordabout 14 hours ago
Not Invented Here.
apiabout 14 hours ago
> every implementation idea rediscovered five times and wrapped in a new README

That's open source since forever, unfortunately.

carloslfuabout 14 hours ago
I agree with the sentiment, but have you seen those videos in which all men say other men are gay? This feels like the same, so much AI paranoia!

I genuinely want to contribute. And hey! I was doing oss this since 2014 so waay before AI was cool.

docheinestagesabout 14 hours ago
It's what happens when you don't do market research.
carloslfuabout 14 hours ago
I'm sorry this makes it seem like I didn't do my research. I did a TON. To fix it I'll add a benchmark/comparison table. Also, I wouldn't call it market research since this is not commercial AT ALL.
genxyabout 14 hours ago
Does a painter check to make sure that a portrait hasn't been painted? What a dismissive comment.
oceanplexianabout 13 hours ago
Half the people on here are using Ollama. No one is doing market research.