Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

73% Positive

Analyzed from 2599 words in the discussion.

Trending Topics

#models#qwen#model#experience#more#different#pro#deepseek#same#using

Discussion (98 Comments)Read Original on HackerNews

sinuhe69about 13 hours ago
The title is misleading. The link led to a pricing page/token plan and not about the new QWen 3.8 model.
Schiendelmanabout 12 hours ago
The submission link is to the twitter announcement. The body just has a different link to pricing.
5701652400about 16 hours ago
in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far. and it is super expensive compare to Deepseek. cannot delegate anything to it, cannot use it real-time low-level tasks either. totally unusable.
3abitonabout 13 hours ago
> in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far.

I have used both Qwen3.6-35B and Qwen3.6-27B locally (both Q8 quantized with llama.cpp). I have also used antirez's quant of DS4-flash. They all performed within the same tier, DS4 being a bit more efficient, but they all gave really good results, mainly used for bash scripting, debugging, python and some C++. I am curious what type of applications/langauges failed with Qwen? One thing to note, the chat templates were "broken" for qwen models and had to debug it, there are already effort on this. Tbh, the same with gemma.

chewzabout 16 hours ago
From my experience Qwen-3.7-Max is above the Opus level but delivers results much faster. Slightly worse then Fable. Way ahead of Deepseek 4 Pro (in speed and overall comprehension) - which is a workhorse on its own. I am using them all with Claude Code mostly.

Qwen-3.7-Plus is quite OK, good for subagent use. Way better then Sonnet.

Qwen-3.8-Max-Preview seems working just fine for me at the moment - I am playing with is right now but too early to say anything. At 10% of regular price it is a steal so far.

gchamonliveabout 14 hours ago
It's useless to talk about models and harnesses without context and method. Depending on how you use the model and what the model is used for, experience may vary drastically. Also, different models with different harnesses require different approaches.

I've been using https://gitlab.com/gabriel.chamon/orisun which is my own simplified methodology, for coding web apps in python and elixir and have been very successful using qwen3.6 27b Q4 locally with help of larger models for architecture, so I get very suspicious when people talk how useless larger models are. They are either using it for a domain that models don't perform well or just not using it right.

taosxabout 13 hours ago
I'm not sure about "useless" but from my experience agentic coding leads to death by a thousand cuts for all projects I've seen so far. Small decisions missed in a codebase that leads to degradation in correctness, reliability and performance. At some point it only takes one engineer to be careless, others skipping PR because they are AI generated...
exceptioneabout 15 hours ago

  > At 10% of regular price it is a steal so far.
What price do you see?

Here standard plan has been discounted to $18.00, from $25.00/month.

ameliusabout 14 hours ago
Can we please include information of what languages we use when making claims like these?

It makes a huge difference if you're writing Javascript/HTML/CSS, Python, or C++/Rust.

Also the application type matters, e.g. user interfaces or scientific computing.

5701652400about 14 hours ago
me: Go, Swift, Kotlin, bash k8s/gcloud

domain: typical web backend tier, mobile apps. not particularly complex, but requires OOP/architecture/system design.

nullbioabout 16 hours ago
If by Opus you mean Opus 4 and not Opus 4.8, then sure.
chewzabout 16 hours ago
> If by Opus you mean Opus 4 and not Opus 4.8, then sure

I meant Opus 4.8 which is rather dumb and ineffective in coding harness, especially with higher thinking levels.

gigatexalabout 15 hours ago
What the difference between your experience and https://news.ycombinator.com/user?id=5701652400? ‘s?

Such diametrically different ones.

big-chungus4about 16 hours ago
Qwen3.7 pro is meh, but 3.7 max is a very good model
Demiurgeabout 14 hours ago
Are these different models or different efforts for thinking (internal back and forth review) using the same model?
2Gkashmiriabout 15 hours ago
Can you tell me more about deepseek?

I paid $2 for deepseek api, put the key in void editor and made a crypto tool in html.

It turned out to be around 67kb. I used sample files in CSV that were a few hundred lines.

It spent around $1.8 in the hour or two or light coding and follow up bugs.

Is it really really this much?

I can't imagine spending a month using it for a day job, it would cost more than the salary so what gives?

I understand the local ai and all that but do cloud providers cost this much?

Earlier I thought "billion tokens" but now not sure

5701652400about 14 hours ago
so Deepseek 4 Pro cannot go on own sessions for too long.

I delegate small-medium tasks: refactors, summaries, research, writing tests + have very good codebase already + extensive history / architecture / docs / linters. so it picks up and does decent small-medium scope work. it is fast, accurate, cheap. does exactly what I want directly and does not waste time nor tokens.

definitely not "implement me complex greenfield project".

k__about 12 hours ago
My 2 weeks with DeepSeek V4:

Pro is ~50% more expensive than Flash.

Both need babysitting.

Plan, split in small tasks, give it docs, types, tests, linter, best practice examples, etc.

Always start a new session when starting a task.

Do regular manual sanity checks, and tell it to find issues in the codebase.

I pay like $1,50 per day for Pro.

5701652400about 9 hours ago
very simlar experience.

I would also add that I run it this way ~12hour a day non-stop. 300M / tokens per day (99.7% cache hit).

h2aichatabout 9 hours ago
I had a similar experience
aduwahabout 15 hours ago
A local AI is not about cost. In fact you will likely pay more for it than with most providers. Just look up the advantages of having access to a technology like this that can be self hosted
ph4rsikalabout 16 hours ago
> Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. D

Anthropic should not have bugged their knowledge distillation attacks.

chewzabout 16 hours ago
> Anthropic should not have bugged their knowledge distillation attacks.

It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons

As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)

RazorBucksICOabout 14 hours ago
Appealing to the Belsheviks for moral authority is, well I will just say an interesting approach. I do not have that much sympathy for Anthropic, but I do not have much sympathy for publishing companies either whose rights to a revenue stream they violated either. Are Chinese AI companies the Robin Hood in this story? Would they be so magnanimous if they had the upper hand? I don’t think so.
beefsackabout 14 hours ago
For those trying to get it to work in OpenCode with a Qwen Cloud Token Plan, this is what worked for me. Note that I've just matched Qwen 3.7 Max for the limits as I don't know exactly what they are.

  "provider": {
    "alibaba-token-plan": {
      "models": {
        "qwen3.8-max-preview": {
          "limit": {
            "context": 1048576,
            "output": 65536
          },
          "modalities": {
            "input": [
              "text"
            ],
            "output": [
              "text"
            ]
          },
          "name": "Qwen3.8 Max Preview"
        }
      }
    }
  }
5701652400about 14 hours ago
also, be very careful which API endpoint and API Token you use. make sure you use right one (obseve your quota is used up. if you hit right endpoint quota used almost immediately). so that you do not accidentally burn API endpoint tokens (they are expensive, can easily hit 200 USD / 3 days which do not count towards your membership "Credits", if you say purchased it with 200 USD signup bonus in Alibaba Cloud)
jxmorris12about 13 hours ago
Why did Qwen stop producing open models? They've gone from building the best open models ~1 year ago to producing like the 10th-best closed models. I don't understand this pivot at all.

Edit: I saw online they do in fact plan to release this openly at some point – x.com/Alibaba_Qwen/status/2078759124914098291

InsideOutSantaabout 13 hours ago
They've announced that they're releasing the weights for a 2.4T model soon:

https://xcancel.com/Alibaba_Qwen/status/2078759124914098291

fragmedeabout 13 hours ago
It's not a pivot, giving away the weights was a marketing strategy that they don't need to keep up with.
Alifatiskabout 15 hours ago
I remember when they released Qwen 3.7 Plus and Max. These models behaved way different from all prior models, it became too verbose. It wrote multiple paragraphs just to answer my prompt instead of the usual concise and direct way responding to me. I didn't like that at all, and I know Gemini also had this behaviour with with the Flash series until I managed to reduce it a bit with personal instructions (in the settings on Gemini website).

I haven't tried Qwen 3.8 Max yet, looking forward to it. My hope is that its way less verbose. Another thing I experience with the Qwen models is that I do not trust their benchmark scores at all. Have anyone played with Qwen 3.8 Max and can share their experience? Which model it come close to? Sonnet 5? GLm-5? DS V4 Pro? Flash? Gemini 3.5 Flash?

lebovicabout 18 hours ago
I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page.

Looks like they're previewing the model only on their subscription plan.

lebovicabout 8 hours ago
(This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan)
trvzabout 16 hours ago
It’s available in the iOS app (or was for me), both logged in and out.
Alifatiskabout 16 hours ago
Is there an iOS app for using Qwen?!
trvzabout 15 hours ago
Yes, but it’s not available in all App Store regions.
siesteabout 15 hours ago
What is a "credit" and how does it translate to tokens for the different models?
xyzsparetimexyzabout 15 hours ago
Its the currency of the future.
ahartmetzabout 14 hours ago
Since Euro and Dollar values are reasonably close, you can call both of them credits, maybe
tclancyabout 13 hours ago
Yes, but credits are money you don’t actually own. So much more convenient, wave of the future and all that.
whynotmaybeabout 14 hours ago
They use it in Babylon5 (supposedly) in 2260!
rhdunnabout 14 hours ago
Does anyone know if they intend on releasing open source/weights variants for 3.8 or whether 3.6 was the last model they are/were doing that for?
rolls-reusabout 14 hours ago
will be releasing weights per their tweet announcing the model https://xcancel.com/Alibaba_Qwen/status/2078759124914098291
antiloperabout 15 hours ago
Does anyone have the privacy policy of their token plan available? Want to check if they retain/train on inputs/outputs.
moffkalastabout 15 hours ago
Lol, lmao even.

Of course they train on literally everything they get their hands on, like everyone else. If you need privacy, that's what local models are for.

adamtaylor_13about 14 hours ago
It's a flippant answer to a real question. Anthropic, OpenAI, and even Grok have "Don't train on my data" knobs.

Whether you trust them is different, but there ARE knobs on other hosted AI companies.

moffkalastabout 12 hours ago
Those knobs don't do anything, don't be silly. It's just optics.
corvabout 14 hours ago
Who is behind this site? Is this another frontend to Alibaba or a reseller in Singapore?
sbinneeabout 15 hours ago
If it offers more than opencode go, the entry plan looks enticing
Advertisement
cakbeslikabout 15 hours ago
Using QWEN models since 2.5. I never used the chat properly but as an API I can say they're quite good, especially when you compare with OpenAI models. Cheaper and almost same level. I will try this now also.
nullbioabout 16 hours ago
I predict that no one will use this and everyone will use Kimi K3.
embedding-shapeabout 16 hours ago
I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.
sunaookamiabout 15 hours ago
Same problem with every chinese model currently, they overthink way too much and take too much tokens and time.
embedding-shapeabout 15 hours ago
More or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :)
EgregiousCubeabout 15 hours ago
A consequence of aggressive distillation?
rurbanabout 14 hours ago
We'll probably use it, but for images. Qwen is still the best for images
rubslopesabout 15 hours ago
Why? Price? If the reason is performance, I've been using non-frontier models for cheap, and they run great for my needs (GLM 5.2, DeepSeek v4 Pro).
jadboxabout 15 hours ago
What's the price difference?
ernsheongabout 16 hours ago
So are locally-runnable models frozen at Qwen 3.6 now :/
worldsaviorabout 16 hours ago
Everyone wanted open models that would challenge Opus and Codex, here, you got it.
ernsheongabout 15 hours ago
We need better coding models that can run on local hardware, i.e. 128GB VRAM or less
zozbot234about 14 hours ago
You can run larger models by offloading to SSD (for weights), it's just slow so people don't do it all that much. But you can get back at least some of that performance by using either MTP (at least for dense models; not effective for sparse MoE models unless you're batching them already and have VASTLY more parallel compute than you'd know what to do with) or batching multiple requests in parallel (note, this hurts throughput for your single sessions but running more sessions in parallel still boosts your total amount of inference. This requires careful management of memory requirements for your context/KV cache, and Qwen models tend to be KV-cache heavy).

Broadly speaking, this ultimately pushes local inference towards a challenging world where you use SSD offload for weights as a matter of course; then smaller requests (or requests sharing the bulk of their context, e.g. subagent swarms) can be batched together and run quickly in aggregate, but running very large contexts will actually limit you to single-session inference and require swapping out even the KV cache itself to some external scratch SSD, further hurting your performance. Then feel free to add wide use of MTP in a probably futile effort to go back to tolerable tok/s numbers.

seanmcdirmidabout 15 hours ago
Queen has that already, although they seem to be moving away from local models unfortunately.
tormehabout 16 hours ago
Is qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious.
ch_smabout 16 hours ago
In my experience, yes. A bit more reliable than gemma for me. I mostly use A3B (35B, mix of experts) though, because it‘s faster, and in the same ballpark intelligence wise as the dense 27B, so it’s the sweetspot for me. I want to try cohere‘s mini code model next, but worried the runtimes aren‘t optimized for that yet.
mark_l_watsonabout 16 hours ago
I found qwen3.6:26b slightly better on my 32G mac mini than the same sized gemma until gemma was updated with better tool support 4 or 5 days ago.

It is like a ping-pong game: the advantage flips back and forth between providers.

regularfryabout 15 hours ago
Worth knowing that Unsloth have just put out another Gemma 4 release from Google's upstream updates which should improve reliability. Bugs in the chat template affecting tool calling and other issues, apparently. https://www.reddit.com/r/unsloth/s/MpC6Hzs4Wj
androiddrewabout 15 hours ago
I have been running 3.6 27b on a dual AMD r9700 setup using Opencode and Matt Pocock's skills workflow for writing Golang CLIs. It's decent, but won't win any awards on code architecture. I guess you can try to AGENTS.md the deficits but I am just exploring its raw Opencode experience right now. Much slower than an API but still 3x times faster than I can read. Tuning it in with a community chat template and a specific penalty for repeats was the sauce needed to get it to work. I can probably start loop daddying it now over the tickets Matt's flow creates.

So yeah, it's the best local model I've seen. I am going to try the Qwopus 3.6 fine tune soon with the same spec and tickets and compare the output of both.

SomeHacker44about 14 hours ago
Would you mind sharing more please? I literally just finished the same set up, with a 9950X CPU and 192G RAM at 4,800 MT/s. I used lemonade with Vulkan and the UD-Q8_x_x model from HF. 256k context. I have about 8G VRAM free, and use the iGPU for my desktop/monitor on Arch. What options do you give llama-cpp or whatever you run please? What other models have you found fit nicely in the 2xR9700? Thanks!!
seanmcdirmidabout 15 hours ago
I actually have long discussions with Gemini about this and have wound up download a bunch of different models for different things. There is no best, just fast but worse, slow but better, agentic or not, reasoning or not great at large contexts, better world knowledge, uncensored, etc…. It’s a bit daunting actually since there isn’t really a one size fits all model that you can just use for everything.
ernsheongabout 15 hours ago
Yes it's between this and Gemma 4 31B which is much slower, but looks like it won't ever get an upgrade. I have to conclude that the MoE variants are unreliable, and MTP sometimes just can't get tricky formatting right.
dofmabout 15 hours ago
The whole series had an upgrade a couple of days ago actually — they have addressed embedded tool calling (and hopefully the MTP formatting stuff though I gave up running the Gemma MTP because it's often slower than not-MTP)

Not tried it yet but I've seen tests that suggest they've properly fixed the tool calling issues.

SwellJoeabout 15 hours ago
I find the 4-bit QAT with MTP to be entirely usable speed on both my boxes (Strix Halo and a desktop with two V620 GPUs, which are slightly faster than the Strix Halo).
cmrdporcupineabout 15 hours ago
For whatever reason prefill (on my DGX Spark) is faster with the Gemma models than Qwen 3.6 models of similar size. On vLLM anyways. Likely just deeply tuned code contributed to vLLM by Google?

vLLM gives me ~7000+ tok/sec with Gemma 4's MoE model. Vs ~6000 tok/sec for Qwen 3.6 MoE.

hnfongabout 15 hours ago
People have been able to run DeepSeek v4 flash with a high spec Mac.
schaeferabout 14 hours ago
I flip flop between qwen 3.6 27b and qwen 3.6 35b 4b active.

But there’s also the quantization of DeepSeek v4 flash called dwarfstar

cmrdporcupineabout 16 hours ago
Gemma4 models are arguably better. Or at least about the same.
atemerevabout 15 hours ago
The best model you can run locally is Kimi K3, as long as you have the hardware. If "what model I can still run on a something resembling something I can put on desktop without separate electricity and cooling water inputs", then it is probably GLM 5.2 (can be run on e.g. Nvidia DGX Station workstation). As long as you have about $100k-$150k.
dofmabout 15 hours ago
Maybe, maybe not. Qwen 3.6 27B is literally just three months old. Hard to predict. Maybe it just wasn't worth making a 3.7, and after all, the 27B release was after the Plus release.
Archit3chabout 14 hours ago
Obligatory "Does it answer security questions?".
cadlernoxabout 14 hours ago
Nice
dluanabout 16 hours ago
waic go brr
vitorgrsabout 15 hours ago
SVG's pelican https://gist.github.com/vitordelucca/521c2d63c9b852c622e7648...

Made on the website, so not sure if on the API there's more thinking options...

esrauchabout 15 hours ago
I feel like the pelican test can't be relevant anymore; the whole point was to to something that wouldn't be in the training set at all and now it is?
rhdunnabout 14 hours ago
A parachuting flamingo? An aardvark driving a bus? It should be easy to randomize the animal and the mode of transport (or vary it with animal playing a sport) to create images not in the training data.
onlyrealcuzzoabout 11 hours ago
How about an animated SVG of a pelican doing the Macarena, profile view, spinning to face the camera on the last beats?
rvzabout 10 hours ago
> I feel like the pelican test can't be relevant anymore;

It never was. The point of this "pelican test" was for performative reasons, or just for attention of the joke.

It is like trying to test whether if an adult elephant could actually climb up a tree and reporting that some elephants are slightly better at doing that than others while also reporting at the same time that they are all bad at tree climbing anyway.

This is an example of testing for the sake of testing. The "pelican test" tests for nothing.

joegibbsabout 15 hours ago
What about an armadillo playing a piano? There are so many potential combinations It would say something if the pelican looked great but the armadillo looked terrible
LatencyKillsabout 15 hours ago
Agree. It was interesting/fun for a bit though.