Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 148 words in the discussion.

Trending Topics

#ascii#byte#detect#encodings#text#usage#clickhouse#simd#optimization#chances

Discussion (5 Comments)Read Original on HackerNews

kstenerud5 days ago
You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop.

Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.

duskwuffabout 3 hours ago
> So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints.

Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.

mort96about 3 hours ago
That, plus markup, be it something XML/HTML-like or something Markdown-like.
vlovich123about 1 hour ago
FWIW I believe the second optimization is at odds with the first. It’s really hard to do this kind of conditional historical check in SIMD.
zX41ZdbWabout 2 hours ago
This is one of the optimizations from ClickHouse - detect ASCII and go a fast path: https://github.com/ClickHouse/ClickHouse/blob/6cde32de1a2463...