Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

50% Positive

Analyzed from 73 words in the discussion.

Trending Topics

#using#runs#llama#tokens#mtp#context#past#few#days#usually

Discussion (2 Comments)Read Original on HackerNews

Pragmataabout 4 hours ago
I've been using it for the past few days, and it runs really well!

I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.

Genuinely very usable, and fully local!

kristianpabout 1 hour ago
Which card are you using? I was getting about 40 with an UD q3 quant with MTP (prediction) enabled and llama.cpp compiled for my compute capability, but was very limited in the context size. I have an 4060 ti 16GB. Wouldn't recommend it as there's a tradeoff between larger context without MTP and about 18 tokens/s.