DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
100% Positive
Analyzed from 90 words in the discussion.
Trending Topics
#highly#paper#effectively#quantized#models#caches#tokens#smaller#title#correct

Discussion (1 Comments)Read Original on HackerNews
Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?
We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).