Hacker News new | ask | show | jobs
by jpease 1 day ago
Why not Mandarin Chinese? Logographic is pretty compact.

Also, this really should be more precise that it’s talking about neo-caveman. Legit caveman no speak English.

10 comments

Actually, when you translate Chinese literally word for word, it sounds a lot like caveman English.

你不去,我也不去

You no go, I also no go

今天这里人很多

Today here person very many

下雨就不去

Fall rain then no go

If you translate English literally (a very analytical language) to a very synthetic language (eg. Latin), it also sounds like caveman spreak
Classical latin is not very order constrained because it embeds so much grammar morphologically; it could aound very much like english or not at all like english, depending on the author's choices.

Vulgar latin does correspond to modern romance languages (and english) much more closely.

Using Chinese could lead to a small amount of gains, but it's nowhere near as drastic as you'd think

Putting aside the question of whether English's larger training corpus would be a dimension of quality: content is "tokenized" before the model sees it. The "~4 English chars/token" rule isn't a strategy to optimize around because in practice, many English words become one token, the same way Chinese words do.

"I love you" and "我爱你" both tokenize down to 3 tokens. So no efficiency gains from a Chinese translation. "Telephone" is just one token, and "电话" (Chinese word for telephone that spans 2 characters) also gets condensed into one token. Sometimes there are minor gains, like "经济" being one token but "economy" being two. But it doesn't end up materializing in the gain you'd hope for by increasing meaning-per-character.

You can play around with OpenAI's tokenizer here. Great for getting particular about how to save tokens in prompt and tool definitions https://platform.openai.com/tokenizer

> Why not Mandarin Chinese? Logographic is pretty compact […]

What do you mean? A lot of us already use Hanzi (汉字) or Kanji.

Or are you asking why non-Chinese speakers do not prompt LLMs in Mandarin?

I cannot tell whether you were really serious but this actually could be an interesting study: same task, same procedure, only the language and writing system is different: English vs Chinese vs Korean vs ...

Since the information density is about the same in all (spoken) languages, there shouldn't be much of a difference. But maybe one writing system is better suited for LLMs.

Heh, deepseek is known to randomly switch to Chinese output for no reason at all. As I was using Claude chat as a poor man’s reviewer, I just pasted it over and Claude interpreted it without skipping a beat and responded in English.
My dad told me he had a German math professor in college that would sometimes have to pause to think while writing out equations on the board and then turn around to a room full of confused students to ask how long he had been speaking German to know how far back in the lecture to restart in English. Again. He was known for doing it. Maybe we can anthropomorphize deepseek similarly??
It would take the typical English speaker many years of dedicated study to learn Mandarin Chinese to a level of fluency required to take advantage of any token savings. "Caveman speak" can be adopted instantly, by anyone (in, presumably, ~any language) with no study or time investment required.
chinese is less token efficient than english
How so? Do you mean in practice when using mostly english trained models?
Even in chinese trained models too. the amount of tokens it requires to communicate the same ideas is higher. it is just less efficient to tokenize the Chinese language because logographic languages are archaic, ridiculous, and inefficient compared to latin languages.
Pretty seen we will all be speaking tokenese anyways
problem is there are thousands of pictograms you have to know with ambiguous sounds

i think korean is the best, just learn the alphabet and you can do pretty crazy compaction by dropping honorifics and abbreviation, you can also sound out foreign words ex. 짱깨 if given the right context LLMs should have no problem understanding