ZH version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
50% Positive
Analyzed from 73 words in the discussion.
Trending Topics
#using#runs#llama#tokens#mtp#context#past#few#days#usually

Discussion (2 Comments)Read Original on HackerNews
I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.
Genuinely very usable, and fully local!