Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

66% Positive

Analyzed from 1714 words in the discussion.

Trending Topics

#llm#embedding#search#data#structured#categories#where#more#query#useful

Discussion (50 Comments)Read Original on HackerNews

Sharlin10 minutes ago
I can’t believe programming is now at the stage where advice like "first have the computer give you totally wrong answers, then just find a function that maps the wrong answers to the correct ones!" is a thing.
kgeist7 minutes ago
A common case I have is when you don't have classifications to begin with. For example, you need to find what users complain about most. I take embeddings of all records, then cluster the embeddings into semantic groups, then ask an LLM to take a random sample from each clustered group and create a classification for that group.

This method is sensitive to the thresholds (what is the maximum distance between embeddings for them to be still considered part of the same semantic group), so I run it all in an agentic loop where an agent tries different thresholds and clustering algorithms until it's satisfied with the result, plus it may deduplicate some groups.

I run it all on self-hosted hardware, so it costs nothing to leave it running for, like, a night, and as a bonus, none of the corporate data leaves the office. I think a rigid set of manually created classifications may not capture all the possible classifications that can exist. Needs a review by a human, though.

alexpotato2 minutes ago
I worked on spam classification for litigation targeting in the early days of CANSPAM [0] enforcement.

We had a similar problem where you can literally millions of email that we were pretty sure came from only a limited set of bad actors.

We first started classifying emails into buckets by From, mailserver relay chains etc as that's all we had to to go on.

Over time, those buckets got linked to spammer signatures and then we narrowed down from there.

Fascinating to see this happening nowadays with LLMs.

Majromaxabout 1 hour ago
> In the notebook, I compute a MiniLM embedding of every real Wayfair classification. I compute the embedding of the fake, hypothetical embedding from the LLM. I then dot product the fake embedding into the real ones to find the most similar. Producing: [the right answer]

Isn't this begging the question that the hallucinated classification will be more selective with respect to the real schema than the query itself? What would the dot product of <E(search query), E(schema)> have given?

Even if that is too vague, smaller LLMs are capable rerankers; return the top N matching true categories and ask for a contextual ordering.

softwaredoug3 minutes ago
Yes what you're describing is a classic way of doing query understanding.

I've found, though, getting it in the language of the vocabulary has generally improved performance.

Further, when searching for "blue shoes" you want to separate the color from the item type. So its useful to have a dumb LLM do this for you. And with the LLM in the loop, its further useful to get it into the language of the taxonomy to improve embedding retrieval accuracy.

There are of course many ways to skin the cat here :)

pu_pe26 minutes ago
Nice trick. Couldn't you embed the query though, compare it to the embedding of the categories, then ship only categories that are close to it in the prompt to a smaller model?
vessenes3 minutes ago
Agreed that you almost certainly can just embed the original with most modern embedding models.
tantalor12 minutes ago
Yeah I had the same question. What's the point of the intermediate step?
HarHarVeryFunny9 minutes ago
Interesting technique, but even if you're getting rid of hallucinations it seems there's still no guarantee of consistent classifications. If you need to do a semantic (embedding) search anyways, then how does this really help?
arjie10 minutes ago
Prompt expansion of input to extra categories makes sense if your embedding isn’t working well. But on its own, why use the LLM at all? I think you could have demonstrated the original step first and then shown that it’s useful.
claudiosf120 minutes ago
Smart trick, but assumes the “dumb” llm is smart enough not to derail into an article about the lives of South American red ants. Obvious exaggeration, the point being outcomes should stay strictly within topic, avoid unrelated bloat and hit the target.
phoghed2 minutes ago
If you use structured outputs they’ll usually stick to the program. Not to completely constrain the categories like TFA was saying, but something like

    { rationale, categories }
Where you don’t really care about the rationale but you’re using it as a pseudo thinking for models that don’t support it.

Luna is surprising capable and cheap, and I haven’t done this type of thing since before GPT 5 so might not be such a useful trick now

smallnix13 minutes ago
Since you map each breadcrumb of the path, how do you deal with differing lengths that would be more appropriate?
piterrroabout 2 hours ago
I would propose the following, query vector store for 10 closest categories based on a query, feed it to an LLM, in the prompt ask it to produce a single digit 0-9 representing the number of the most appropriate choice. Use plain text prompt, dont inflate token count with JSON. There you go, you just drastically reduced the output pricing.

Additionally you could experiment with a reranker instead of an LLM or after reranking take top-3 results and then feed to LLM as input in order to reduce input token costs.

jddjabout 1 hour ago
Or press 9 to hear these options again
virgil_disgr4ce14 minutes ago
Your call is important to us. Please listen carefully, as our menu options have changed.
cesargstnabout 2 hours ago
good this yeah
eka1about 2 hours ago
Did you validate this by running a A/B test? Main question is were you able to classify back into your known categories correctly all the time, or did the errors compound from the llm hallucination plus embedding search
softwaredougabout 2 hours ago
Using a Nano model, a tad worse than shipping a vocabulary to a larger OpenAI model. (And it’s an huge improvement on not classifying the queries at all).

But no classification is perfect. In search in particular, you will also want to have places for manual intervention for high priority queries.

iandanforth12 minutes ago
No? This is just giving up and hoping.
motoxpro7 minutes ago
Is there a solution you are using to solve this that is more accurate and cost effective? I'm working through it now so would be curious
runarberg3 minutes ago
Is scraping and putting this in a structured format too inaccurate or expensive?
chrisjj10 minutes ago
[delayed]
Advertisement
ipsodabout 1 hour ago
Just this week I tried doing something similar with a nasty vibe-coded codebase I was trying to organize. I had Gemini Flash 3.6 classify each function/method in a similar way, giving a few plausible classifications for each (one agent per method).

It didn't end up being very useful - I ran a comparison where I just had a bigger agent do the organization in a more straightforward way, and that had better results.

I did find that Flash 3.6 High was >9x faster than Luna xhigh for this task, and got very similar results, though.

sheepscreekabout 1 hour ago
I’ve read a few different accounts, including OpenAI’s own admission, that Terra Medium or higher will likely produce better results than Luna xhigh and cost about the same or less.
amitpoonia19xyz37 minutes ago
This is basically HyDE (Hypothetical Document Embeddings), no? I had tried this approach in the past, worked with limited success.
memjay26 minutes ago
We have this running in production. Can get pretty expensive and slow. We are trying to replace this with cheaper and faster methods that don’t hammer our LLM and elastic search endpoints as much.
thatjoeoverthrabout 2 hours ago
Smart! I've done the same trick for resolving extracted intents to selection.

But if accuracy matters, you can't rely on embedding sort to get a closet match. With a real test set they usually don't hold up under scrutiny.

Everything in AI is like this. You get an idea, try it once or twice, "LGTM" and you ship. Then it never survives contact reality.

Embedding sort gives you a better shortlist than the whole list, but you will probably want a heavier model to vet candidates.

ashu1461about 1 hour ago
Had stumbled on this library in the past

https://github.com/aurelio-labs/semantic-router

I guess it is based on the same fundamentals as well.

estetlinusabout 2 hours ago
I was in a project where we sent the whole taxonomy every request, 40k tokens + one article, ”plz classify”. This was before structured outputs. It was extremely expensive and still hallucinated. Good ol’ days.
ameliusabout 2 hours ago
Can anyone explain why LLMs are so bad at finding products (their webpages) with given specifications?

You'd think they would have solved it by now.

simonwabout 1 hour ago
LLMs aren't architected to handle filter-style comprehensive search without setting them up with additional tools.

Asking an LLM for a list of every county in the USA for example, or every county with a population of more than 100,000 people.

Even if those county names and their populations are mixed up in their weights, the nature of next-token-prediction does not lend them to effectively answering comprehensive, detailed questions like that.

An agent system build on top of an LLM can do it, if it has access to tools which can help access eg a table of counties and then filter them with SQL or Pandas or similar.

amelius38 minutes ago
Yes, I was assuming they'd use external tools. Using only the raw LLM doesn't sound like a good strategy.

Considering that agents are not a new concept, why isn't this a solved problem by now?

ACCount3721 minutes ago
Agents are a very new concept.

We've got the early LLM-based AI agents in 2023, and it only became a popular, mainstream thing in 2025 - with Claude Code.

braiampabout 1 hour ago
Because that's structured data and structured data is usually hidden away from users _and_ machines. Product rarely want to be honest, unless it's B2B in a very competitive market (and even then!). So, yeah, it's not that they are bad, it's that there are few good sources of information.

(Lets ignore for now that no one seems to agree to what should be the spec sheets)

ashu1461about 1 hour ago
With agentic commerce protocol / unified commerce protocol open ai and gemini are trying to solve this problem.

The idea is to make structured queries using these protocols which can be used to fetch top products matching the user needs instead of just relying on semantic search.

https://developers.openai.com/commerce/specs/file-upload/pro...

Zigurdabout 1 hour ago
It gets worse: shopping agents are hostile adversaries to Amazon unless they're paying Amazon and they've agreed to be friendly agents. No agent that won't betray you to an Amazon pricing strategy is going to be allowed access to Amazon structured data. They might even be fed poisoned data to discredit them.

But you'll be amazed by the abundance.

ameliusabout 1 hour ago
An LLM can read websites, right? And turn them into structured data.
ashu1461about 1 hour ago
It can do that on run time, but it does not store data like that. The data is typically stored as embeddings in which it is hard to query data in a structured form. Example give me all products whose price is less than 200$ vs suggest me products for my spouse's birthday.
_fluxabout 1 hour ago
Amazon Rufus has been mildly successful for me. I think the failures I've experienced with it are mostly because the product I'm looking for doesn't exist in the catalog.
pydryabout 1 hour ago
It's been an absolute fucking disaster for me. It hallucinates endlessly and its searches are terrible. It even managed to confidently gaslight me about there being a VAT invoice available for a specific product.

I noticed yesterday when browsing on mobile that there used to be a box where I could search reviews and it got swapped with a Rufus box. I guess somebody needs to juice their engagement numbers for an investor briefing.

honestly, Amazon doesnt even need AI it just needs a better UI, more metadata for its products and to make reviews less scammy.

sgcabout 1 hour ago
I asked a question once and now there is a effing alexa for shopping toolbar that takes a quarter of the screen that will not go away no matter how many times I close it, and the space remains taken even if I adblock it. Absolutely hostile implementation. I have words for this I cannot type out.
wslhabout 1 hour ago
Because the data, in general, is not included in the LLM model and it needs to search/browse for external information. It cannot look indefinitely so it get the top results from lists, not "evrything".
Colegnoabout 2 hours ago
Isn't search engines quicker than calling a LLM ? It might have a huge impact between a 20ms search engine call and a 2s LLM call for the end user.
quixoticaxolotlabout 1 hour ago
They are already solving the problem with search engines, they're just using an LLM as a first pass to create better embeddings to run a similarity match on first. The difference in latency is likely made up for in accuracy.
fastballabout 2 hours ago
A 2s LLM call is pretty slow.
gadflyinyoureyeabout 1 hour ago
Try using Digital Ocean. Minutes spent on inference.
sergiotapia39 minutes ago
This is a really great trick, woah!
apwheeleabout 1 hour ago
This is another riff on not embedding a full document, but doing a summarization of the document and embedding the summary for RAG. Nice usecase for high cardinality data!
VladVladikoffabout 2 hours ago
Eh, maybe you should keep both paths. When LLMs eventually crawl the site to feed back to agentic shoppers, maybe they logically follow the more truncated less decorated path.
Advertisement
einpoklumabout 1 hour ago
In the past, people would post advice on how to do something clever and useful yourself. Now, people post suggestions on how to talk out the side of their mouth to coax ther magic-8-ball slop generator to say something useful.