Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

80% Positive

Analyzed from 623 words in the discussion.

Trending Topics

#caveman#english#token#why#chinese#should#nerd#mandarin#pretty#more

Discussion (24 Comments)Read Original on HackerNews

DougN7about 1 hour ago
I don’t know about the rest of you, but the bulk of my token usage is not what I type in - it’s that data the AI is processing for me. I’m shocked they even got an 8% reduction.
jpease27 minutes ago
Why not Mandarin Chinese? Logographic is pretty compact.

Also, this really should be more precise that it’s talking about neo-caveman. Legit caveman no speak English.

olalonde9 minutes ago
Actually, when you translate Chinese literally word for word, it sounds a lot like caveman English.

你不去,我也不去

You no go, I also no go

今天这里人很多

Today here person very many

下雨就不去

Fall rain then no go

guessmyname19 minutes ago
> Why not Mandarin Chinese? Logographic is pretty compact […]

What do you mean? A lot of us already use Hanzi (汉字) or Kanji.

Or are you asking why non-Chinese speakers do not prompt LLMs in Mandarin?

sparky_z18 minutes ago
It would take the typical English speaker many years of dedicated study to learn Mandarin Chinese to a level of fluency required to take advantage of any token savings. "Caveman speak" can be adopted instantly, by anyone (in, presumably, ~any language) with no study or time investment required.
ButlerianJihad17 minutes ago
zuzululu22 minutes ago
problem is there are thousands of pictograms you have to know with ambiguous sounds

i think korean is the best, just learn the alphabet and you can do pretty crazy compaction by dropping honorifics and abbreviation, you can also sound out foreign words ex. ㅉㄲ if given the right context LLMs should have no problem understanding

NitpickLawyer24 minutes ago
I think it's funny this works somewhat, but that's not what the goal is, IMO. The goal is to have the model do this internally in its "thinking" stage. On the very rare occasions in the past where gpt5 leaked its true internal thinking, the "CoT" was itself kinda similar. Instead of the open source "so the user wants me to... but wait... maybe I should... blahblah...", GPT5 was using internal traces like "try x.. no.. try y... no.. from x yes then z yes...". That's probably more token saving, or faster responses, if you can get the model to remain accurate.
pineappletooth_about 1 hour ago
Well 8%-10% saving without measured quality degradation is not nothing, specially now that newer models seem to use more tokens than previous ones.

Also it was tested on reasoning low, i'd have liked to have them tested on higher reasoning levels.

mjevansabout 1 hour ago
Ask, what is the function of grammar in language? I would argue two major functions. Forcing an encoder to marshal ideas into a coherent structure. Preserving the integrity of that structure when communicated and decoded.

Grammar sounds off? Was the message correctly understood; is there enough left to apply error correction and regenerate the intended message?

mmastracabout 1 hour ago
I suspect this means that there's just a ~10% inefficiency in token to information mapping.
igor_nast3 days ago
As the test shows the real gain is somewhat marginal to most people, but still some ppl claim a wide skills portfolio is required to move faster or save tokens. Mixed feelings, I personally use superpowers as my basic setup - and only them.

What are your thoughts on it?

glitchcabout 1 hour ago
I use my local GPU. Yes, I realize it's not exactly a superpower, but it's technology that's pretty darn close to one.
abofhabout 1 hour ago
If you just want a smaller vocabulary, use French? If your goal is to communicate to an LLM, maybe saving tokens isn't the all in win, unless you like reading assert gronkHitThing(true)
kardosabout 1 hour ago
French text is somewhere around 10-30% longer than the corresponding English text. I would guess much of what you save on smaller vocabulary is lost on the lengthened text.
ViscountPenguinabout 1 hour ago
What would matter more is the token length, depends how well your tokenizer was trained on french I guess.
ButlerianJihad42 minutes ago
https://groups.google.com/g/alt.nerd.obsessive/c/EGuKIN_4dME...

  alt.nerd.obsessive FAQ v1.4

  In the episode where they were filming the Radioactive Man movie [1995], the comic book store guy tells Bart that he can find out the star of the RM movie. He promptly posts a message to alt.nerd.obsessive which states "Need know star RM pic". This information is relayed through the nerd world until it reaches a nerd hiding under the table at a meeting of movie moguls casting the film. He relays back the answer immediately (Rainier Wolfcastle).
https://en.wikipedia.org/wiki/Radioactive_Man_(The_Simpsons_...
sh34rabout 1 hour ago
gpt 4 moto razr wen?
altmanaltmanabout 1 hour ago
Yes all the caveman spoke English like its freaking Flinstone. Like wtf is the caveman dilect, just english words with some randomly skipped? Why not try with actual other human languages and see if there is a token benefit in savings
gruntled-worker32 minutes ago
Me mechanic not speak English. But he know what me mean when me say "car no go", and we best friends. So me think: why waste time, say lot word when few word do trick?
altmanaltman27 minutes ago
To see world
strictneinabout 1 hour ago
Sounds like some good research. You should pursue it.
altmanaltman26 minutes ago
Someone definitely should
inasenseabout 1 hour ago
lol - so easy a caveman can do it.