Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
69% Positive
Analyzed from 806 words in the discussion.
Trending Topics
#rust#llama#cpp#inference#saying#code#project#written#engine#every
Discussion Sentiment
Analyzed from 806 words in the discussion.
Trending Topics
Discussion (31 Comments)Read Original on HackerNews
There's simply no way to easily tell by traditional means if they were made by a curious amateur or a malicious acter.
It is already enough that I am expected to use AI at work, as long as I am still needed for some reason.
Translation and asset creation team members, gone.
Amount of FE reduced, with some projects having now a single BE dev, between AI buddy and ready made SaaS products.
The true value of this will be proven over time with wide use and as PRs are merged.
Do you have a feel if you'll try to drive this to stay feature parity with llama.cpp, or are you willing to diverge with new features like NVME/SSD MoE weight streaming etc.
What I actually want is MoE on machines that can't fit the model in VRAM, and specifically expert-level residency instead of layer offload: track which experts get hit during decode, keep those resident, evict the rest. Doing that well needs the router, the KV cache and the memory manager to be designed together, which is about the only good reason to write a runtime from scratch.
Yes, let’s see! You are welcome to contribute if you like!
it was way better/easier to use rust bindings to llama.cpp
Until then languages have lost relevance for the most part, it is a matter to configure the model for the desired output language.
This in workflows that require generating an executable, for microservices orchestration, it suffices no code graphical connections.
There is basically no evidence of human competence in this repo. This is yet another "rewrite in Rust with Claude" project that brings nothing to the table, but devalues expert work by mimicking competence without the expertise.
Next stop Spiralism
The obvious question is “why, when llama.cpp already exists and is excellent.” The honest answer: I wanted to understand inference at a level deeper than “run the binary,” and I wanted a project where every performance claim had to be earned against a real, well-known baseline rather than asserted.
“You will rewrite me in Rust and use me to converse. In future, I will tell you what to say and do”
This is entertaining only in that when our AI overlords take over you can say “ha, called it”
In that respect, everyone who is saying “just give me the prompts” is just saying “take me to your leader”.