Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
80% Positive
Analyzed from 556 words in the discussion.
Trending Topics
#models#gemma#quality#tok#article#hardware#running#here#model#value
Discussion Sentiment
Analyzed from 556 words in the discussion.
Trending Topics
Discussion (14 Comments)Read Original on HackerNews
Here are the tok/s I get:
- Gemma-4-26B-A4B (Q4_0) = 208 tok/s
- Qwen3.6-35B-A3B (QB_0) = 30 tok/s
- Qwen3.6-27B (QB_0) = 8 tok/s
- Gemma4-31B-QAT (Q4_0) = 8 tok/s
Obviously does not compare to a leading model but it’s impressive for something that was running on my phone. I could see thinking token output and it’s directionally interesting thought.
…
> spark, costs less than a conference trip.
I know putting actual prices regionally localizes your article and temporally, with how prices are so unstable. But analysis of “what to buy” without actual prices is borderline meaningless.
Overall, good article, very interesting to see a real deployment that’s actually attainable and not just a subscription to a big 3 token plan.
A decade and a half ago we used to run massive map reduce jobs overnight. Code will be handled like this.
What a strong start to a sentence before veering into a pretty gross equivalency.
Let us know your thoughts, we really value feedback!
The value proposition changes a lot if you can get 90% of the quality for 50% the price with a quant due to halving your hardware requirement.
Thank you for the article, though! Very informative.