HI version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
71% Positive
Analyzed from 362 words in the discussion.
Trending Topics
#search#benchmark#engines#queries#seems#engine#github#benchmarks#results#trivial

Discussion (8 Comments)Read Original on HackerNews
what if every engine returns garbage, or on the other hand, handles them too well?
building a benchmark like this in a genuinely fair way seems extremely hard to me. I’m very curious about the details, of course within what you can share.
The queries from what I can tell are not trivial. The actual github repo of the benchmark has a judgement/query browser where you can inspect the different query streams: https://keenableai.github.io/needle/
We developed this live benchmark with daily/hourly sampled fresh queries matching real agentic search traffic to estimate actual search performance of different AI search providers.
What do you foresee in the future releases and improvements to this?