Back to News
Advertisement
SSilenN 4 days ago 23 commentsRead Article on github.com

RU version is available. Content is displayed in original English for accuracy.

Hi HN, we built world-model-optimizer, an open source tool to continually improve a specialized model for an agent.

It does this by simulating production tool responses through text world modeling (similar to QwenAgentWorld, summary here https://x.com/silennai/status/2073887455884058814).

We can then use this to train a router for frontier, OS, and local models (use defaults or pick which ones to optimize against).

wmo ingests agent traces, builds the simulation, embeds the traces, runs different models you choose against the simulation scenarios, and then uses a KNN for model selection (similar to https://arxiv.org/abs/2505.19797).

- Cache aware: cache is taken into account for the effective price in routing.

- Confidence gated: we don't deviate from the best fit model when paired evidence over retrieved neighbors is below 0.5 standard errors or on queries unlike anything in the fit set.

- Optimize for cost or quality: train a balanced, cost max, or quality max router.

Usage

`wmo build` creates the simulation (or add your own benchmark)

`wmo optimize` tunes the router

`wmo serve` starts the server and can run everything fully locally. The simulation and router can update over time as more agent traces are gathered and new models are added.

Router results vs Fable

- RouterBench: -66.5% cost, -1.7% performance, -24.7% latency p50. 77.5% of traffic to Sonnet 5, 16.1% Fable 5.

- TauBench: -44.5% cost, +6.3% performance, -20% latency. 83% to Opus 5, 17% to Kimi-K2.6 (over K3).

- Terminal Bench 2: -64% cost, +8% performance, -50.6% latency. Sonnet 5 is fully along the pareto front. Training a specialized router per task isn't cheap. In sparse data regimes the value can be "here's the best model".

We're working on sample effiient continual learning for agent specific models at experientiallabs.ai"

Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 400 words in the discussion.

Trending Topics

#models#model#more#local#tuned#https#com#frontier#cool#thanks

Discussion (23 Comments)Read Original on HackerNews

adrianco4 days ago
Local models need to be tuned to work well so this looks useful. Seems to be for general purpose model serving. I’ve been using https://github.com/adrianco/retort to run experiments for coding models across 13 different programming languages to see which frontier and local models work.
SilenN4 days ago
That's cool, thanks for sharing!
anshad2u2 days ago
Interesting approach. What does the cold-start phase look like for a new agent? How many traces or runs do you typically need before the router has enough signal to safely offload tasks from the frontier model??
jack_pp4 days ago
Not sure I get it. The model you're improving is local? If so how do you even calculate cost compared to an API
SilenN4 days ago
Open source models.

wmo routes requests between frontier models and open source models that continuously train using Tinker. As the smaller models improve, more traffic gets routed to them.

Calculating cost is just tokens in/out.

handfuloflight8 minutes ago
What are the costs to train and use the Tinker models?
surround4 days ago
The title is misleading. This is model routing, not distillation.
SilenN4 days ago
Fixed formatting which will help with readability. We do routing, distillation, and token compaction.
digitaltrees4 days ago
Cool project
SilenN4 days ago
Thanks :)
yiyingzhang4 days ago
Cool idea! How do you guarantee privacy?
SilenN4 days ago
It's open source!

We do have a platform we'll be launching as well to manage training + serving for you which will require more diligent privacy guarantees.

rglover4 days ago
Excited to play with this.
SilenN4 days ago
Let me know if you have any questions!
Art96814 days ago
The absolute best way to prove this works is by releasing a model that was fine-tuned with this method and then showing benchmarks depicting the improvement delta between the base model and the fine tuned one.

The work is not done. Then release it to the masses and wait a few days for the actual real world anecdotes.

Until then, this is noise.

SilenN4 days ago
Valid criticism. Happy to answer any qs. We're still working on solidfying results.
dang4 days ago
Ok, I think it is in your interest to wait until you have more to show, and we'll be happy to help you with reposting it once it's ready.

Waitlists are against the Show HN rules (https://news.ycombinator.com/showhn.html), and you're likely to get a lot of community pushback if you post before there's enough substance for users to sink their teeth into.

Edit: we eventually got a more substantive writeup from OP so I moved that text to the top and re-upped this thread.

SilenN4 days ago
Thanks for the heads up, removed mention!
irishcoffee4 days ago
Benchmarks are the ultimate consolidation of halnons razor.
teravor4 days ago
[flagged]
dang4 days ago
"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something."

https://news.ycombinator.com/newsguidelines.html

Reubend4 days ago
Yeah, this is just slop. No benchmarks, no concrete case studies, just some vibecoded "platform" to finetune models on your own traces.

Which is an idea that has some value, but also some weaknesses. And this implementation of it isn't forthcoming with that concept. You have to really dig in to understand what they're even talking about.

SilenN4 days ago
Happy to answer any qs.