Advertisement
Advertisement
β‘ Community Insights
Discussion Sentiment
33% Positive
Analyzed from 175 words in the discussion.
Trending Topics
#simd#more#data#cache#stores#bad#ops#throughput#cpu#load
Discussion Sentiment
Analyzed from 175 words in the discussion.
Trending Topics
Discussion (6 Comments)Read Original on HackerNews
Also, I see no mention of alignment in the post. I understand x86/AVX2 likes your load/stores to be aligned, even if it technically allows unaligned access.
So, for 32B loads/stores and 64B cachelines, it's 1.5x more L1 cache ops (as half of the ops will cross a cacheline); perhaps bad if you're L1-cache-throughput-bound, but less so if you're at L2+ as the extra work sits in L1.