Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 630 words in the discussion.

Trending Topics

#model#https#more#size#laguna#flash#poolside#models#things#deepseek

Discussion (27 Comments)Read Original on HackerNews

Lwerewolf•about 1 hour ago
Testing it now. At the very least, competitive with DS4-Flash indeed. On my small (and per Sol's words, _very_ semantically dense) C test codebase, it found things that only gpt-5.2 managed to find back in the day, but also made a stupidly incorrect initial observation that a memfd_create()/mmap was used for IPC (funnily enough - sol missed that as well in its review, until I pointed it out). Re: the claims vs deepseek v4 - both flash and pro are expected to get a "general availability" release very soon (i.e. well-"post-trained"), so things can change in a... well, flash, as per usual in the current environment.

Anyways, keep 'em coming.

ilc•about 1 hour ago
What harness/quant did you use for testing?
Lwerewolf•about 1 hour ago
nvfp4 mlx, literally barebones pi.
mft_•about 1 hour ago
Looks impressive, and this size fits achievable home hardware.

That said, if someone would kindly quantise this down for the 64GB paupers, that would be appreciated. (I know there’s likely degradation, but some people reported good results with a 2 bit version of Qwen 3.5 122B, and this is starting from a higher point. Would be interesting to try, at least.)

Edit: someone in the process of doing so: https://huggingface.co/vcruz305/Laguna-S-2.1-GGUF

yogeshp•10 minutes ago
They have also published smaller 33B model called Laguna XS 2.1, its Q4 gguf is 20GB.

https://huggingface.co/poolside/Laguna-XS-2.1-GGUF/tree/main

verdverm•19 minutes ago
The tool I've been using, llm-compressor, can quant models that do not fit in memory (use the sequential pipeline)

https://github.com/vllm-project/llm-compressor

my setup to help you on your way: https://github.com/verdverm/quantr

Though it seems these will not be needed as Poolside has published quants & dflash with their models.

mchusma•about 1 hour ago
Incredible. This is definitely the launch of the day. Just crushing Google's releases.

The pricing here is incredible. This is the first US release that's competitive with DeepSeek V4 Flash. Very excited about this.

river_otter•about 1 hour ago
Hey, this model is not a joke! Exciting, we already got a usable PR of work out of it.

https://github.com/mozilla-ai/otari/pull/348

kamranjon•about 1 hour ago
Whoa whoa whoa, 118b params, 8b active MOE, long context reasoning, open weights - music to my ears. Hadn't heard of this lab before but I am very excited, will definitely try this out tomorrow - this is a real sweet spot I think in terms of model size and performance.
svclaws•42 minutes ago
If the numbers are legitimate then our prayers have been heard
Iolaum•about 2 hours ago
Model Looks amazing!

Even more important, subjectively, is that this model will run very well on Strix Halo (e.g. Framework Desktop), DGX Spark kinds of devices. Looking forward to Unsloth dynamic mtp quants.

P.S. Looking at the HF release they already offer Q4_K_M and DFlash drafter for speculative decoding!

verdverm•18 minutes ago
I hope all models going forward come with a dflash drafter so we don't have to train one up separately.
SwellJoe•about 1 hour ago
This is exactly the kind of model that's been needed in the middle. Realistically self-hosted, Good Enough intelligence, MoE so it's fast on limited bandwidth systems like Strix Halo and DGX Spark.

For a while there's been nothing to run on my Strix Halo that's notably better than what I can run on my dual 32GB GPU desktop (Gemma 4 or Qwen 3.6 dense models), but this seems likely to be the step up in size that actually works better than those.

carimura•10 minutes ago
Congrats Poolside team!!
river_otter•about 1 hour ago
I love this. Is it possible to give a feel of how this stacks up to the good old Opus 4.5 in coding quality? For me that was the turning point where agentic coding in Claude Code etc became usable. Have we hit that threshold?
megavon•about 1 hour ago
Having played with it for like 3 hours now....I'm probably moving from CC to this
fingerprinter•3 minutes ago
One hour in, no more Codex for me. This thing rips.
river_otter•about 1 hour ago
I am about 1 hour into using it with pi.dev. Do you have thinking on high? It is doing good but at one point i had to stop it and say 'you're overthinking this' haha
megavon•18 minutes ago
Yes full send mode on thinking. I have moved on from watching my agents and I don't really care how it thinks. I look at the end result and so far this thing has been blowing me away. No way this is as good as it is this small and fast. Outside Fable, this might be the best thing I've ever used.
kamranjon•about 1 hour ago
What quant are you using?
megavon•about 2 hours ago
This is INSANE. How did they do this?
eisokant•about 2 hours ago
"What we've done in this model is not necessarily add more intelligence, but improve the behaviors that lead to a more capable model: more verification, less taking things for granted, not declaring victory early, and being more persistent.”

+

https://poolside.ai/assets/laguna/laguna-m1-xs2-technical-re...

Lwerewolf•about 1 hour ago
Almost like a built-in heavyweight harness.
kamranjon•about 1 hour ago
"It went from the start of training to launch in under nine weeks..."

This is pretty impressive.

tosh•about 2 hours ago
open weights and

similar performance to deepseek v4, inkling at size of nemotron 3 super (!)

Advertisement
iraldir•about 2 hours ago
Amazing model at this size if true, that's quite crazy!