Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 90 words in the discussion.

Trending Topics

#highly#paper#effectively#quantized#models#caches#tokens#smaller#title#correct

Discussion (1 Comments)Read Original on HackerNews

DiabloD3•about 1 hour ago
The title of the paper is correct. The paper does not seem to actually get to the point in a generic way, but hyperfocuses on, effectively, one type of error compensation.

Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?

We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).