Back to News
Advertisement
ggioscarab about 12 hours ago 23 commentsRead Article on github.com

DE version is available. Content is displayed in original English for accuracy.

I love research and development, you may have heard of me because of PJON (Padded Jittering Operative Network). It is a network protocol I started developing in 2010, which was recently implemented in silicon by the ETH Zurich university thanks to the research of Pius Sieber.

I am excited to share with you TERMy, a terminal assistant built on top of the NPC-Forge framework. Unlike everything else being built today, TERMy does not use embeddings, machine-learning or LLMs. It runs on the CPU (even on a Raspberry Pi Zero) both in the terminal or client-side in a browser tab and responds in milliseconds. It is a cynical but very knowledgeable Linux terminal assistant that translates your natural language into shell commands without relying on a single artificial neuron.

I had a chance to focus for 2 months on my personal projects since early July, during the strange times of AI price hikes and the end of subsidized tokenmaxing. I was curious to see if I could develop from scratch a terminal assistant capable of handling simple natural language requests. I have a bad memory and got used to ask to copilot "activate the virtual environment" or similar trivial operations spending a non negligible sum every month. I started thinking, maybe I can do something to make my workflow more efficient? Do I really need trillions of parameters to accomplish those tasks?

How it Works

When you type a prompt, it goes through a lightweight NLU pipeline written in ~1000 lines of Python that implement the following steps:

1. Strip expletives, interjections, encouraging, discouraging and thanking words (remove noise)

2. Sentiment analysis

3. Exact Match (very fast)

4. Template Match (slower)

5. Probabilistic Match (even slower)

Step 5 relies on:

1. IDF (Inverse Document Frequency) to identify rare words.

2. BOW (Bag Of Words) to accommodate word inversions.

3. IDF weighted Levenshtein to safely handle typos.

Permission gating is hardcoded into the dataset and enforced for all potentially destructive commands, so it's inherently safer than letting an unpredictable LLM run wild on your machine.

- TERMy in operation: https://www.youtube.com/watch?v=qeIp0xePLBg

- Variance and typo tolerance: https://www.youtube.com/watch?v=tQvGDk6fkk0

- Copilot integration: https://www.youtube.com/watch?v=Wzzouhq2a8A

- Advanced features: https://www.youtube.com/watch?v=qeIp0xePLBg

- Source Code: https://github.com/gioblu/NPC-Forge

Advertisement

⚡ Community Insights

Discussion Sentiment

67% Positive

Analyzed from 653 words in the discussion.

Trending Topics

#termy#output#deterministic#llm#https#test#txt#tool#dataset#different

Discussion (23 Comments)Read Original on HackerNews

mbil6 minutes ago
It's kind of antithetical to the tool's deterministic positioning, but have you considered making TERMy leverage an LLM for unseen or low-confidence queries, and then generate the config and update itself to make future similar queries deterministic?
vegnusabout 3 hours ago
If you could get Termy to code, you'd be a rich man
gioscarababout 1 hour ago
Thank you very much for the link.

WOW! With that dataset the capabilities of TERMy could be vastly extended!

Thank you.

gioscarababout 3 hours ago
Hi, I am the creator, feel free to ask any questions :)

What do you think about it?

gurjeetabout 3 hours ago
I haven't evaluated it yet, but I love the fact that the output is (at least claimed to be) deterministic. I can't trust an LLM to do the right thing after I deploy it to production, because their output is non-deterministic by design.

TERMy (or is it the NPC-forge) seems to be worth a try.

piterrroabout 1 hour ago
You can get determinostic output (mostly) by setting the temperature to zero. Using couple of other tricks you can get close to 100% of determinism with LLMs.
jdiffabout 1 hour ago
That's reproducible, I wouldn't call it deterministic. Small, semantically meaningless changes in the input can still result in wildly different output.
kouteiheikaabout 2 hours ago
> because their output is non-deterministic by design.

It isn't. At least not by design, even though in practice it often can be. If you do greedy decoding (or use a preset seed) and deterministically compute everything (e.g. only use integer math) then it will be 100% always deterministic.

kennywinkerabout 1 hour ago
That’s true, but not true-true. Sure, every time you prompt “what is the weather in kansas” you’ll get the same output, but if you prompt “what is the weather in kansas right now” you’ll get a different output, and then “what is the weather in kansas today” gets a different output. Language being language, there are infinite ways to say things, so there are infinite variations in what the llm can output in response to very similar prompts.

This tool has a finite amount of outputs for an infinite amount of inputs. Which is different from an llm based tool.

kouteiheikaabout 2 hours ago
> Models like ornith:9b, mistral:7b or cogito:14b can get the job done sometimes, but they are not fast and reliable enough for general use, specially if you have only 4GB of VRAM.

Have you considered/tried using a model that's, well, more appropriate size-wise for an use case like this? These are relatively big. Something like FunctionGemma [1] finetuned for a given set of tasks would be a lot more speedy.

[1] https://blog.google/innovation-and-ai/technology/developers-...

coder543about 2 hours ago
FunctionGemma never worked well for me (without fine tuning). Liquid has released 230M and 350M models that work far, far better in my testing: https://huggingface.co/LiquidAI/LFM2.5-230M

I really look forward to a hypothetical LFM3-230M, because LFM2.5-230M is so close to being usable, while FunctionGemma is miles away from being usable.

But, yes, still tangential to TERMy.

kennywinkerabout 1 hour ago
https://github.com/ThorOdinson246/whatisit-nl2sh uses a finetune of Qwen2.5-Coder-1.5B-Instruct. It works pretty well, tho it will misunderstand things from time to time
gioscarababout 2 hours ago
I tried functiongemma, it is for sure faster than those models, the problem is that is not reliable enough for a terminal assistant. I would say that no LLM is good for a terminal assistant, if you take into account the operational cost and the risk of damage. Even if it fails only 1 time out of 10 becomes useless. That's why I developed FlintParser!
registereduser1about 3 hours ago
Cool project! How does it differ from warp terminals ai mode where you can ask it questions and it responds back
gioscarababout 3 hours ago
Warp uses LLMs so it is slow and prone to hallucination. Using very colloquial terms TERMy is more or less a calculator that knows english :) so it can run on your CPU and respond instantly! The difference is that it can only answer predetermined responses (with optional arguments) this makes it useless if you need to generate text, but makes it safe and predictable for a use case like a terminal assistant.
utopiahabout 2 hours ago
What dataset does step 5 rely on? Is it from your own terminal history, man pages, scrapped dataset from e.g. StackOverflow, sth else?
gioscarababout 2 hours ago
The dataset is here: https://github.com/gioblu/NPC-Forge/tree/main/npcs/termy/dat...

I hope the community will help me to enhance it :) it is just a proof of concept for now

mpalmerabout 3 hours ago
At first blush, it is a really persuasive compromise between full-on LLM inference and boring old fuzzy history search!

I really like it, this flavor of specialization gives the user a win on privacy and speed. Seems like the right idea for such a tool.

indigodaddyabout 2 hours ago
So is this kind of like a super-powered tealdeer ?
gioscarababout 1 hour ago
tealdeer just shows you a cheatsheet, termy can effectively take a prompt and execute a command, example:

$ termy create file test.txt and write Hello

TERMy | template match | Confidence: 100.00%

Thinking: Ok, I am asked to create the file test.txt.

echo 'Hello' > 'test.txt' && termy_set_context 'active_file' 'test.txt'

Description: Writes Hello in file test.txt.

Response: Affirmative

Now that I think about it, I should let TERMy use tldr...