| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by xvilka 1138 days ago
	Same with Chinese language, thus lexing and parsing requires knowing many more words than in languages with spaces between words.

2 comments

zdragnar 1138 days ago

As an English native speaker who learned Mandarin, I really didn't find the lack of spaces harmful to learning the language.

Since each character represents a syllable, rather than a specific sound, and the written language is essentially not phonetic, reading the characters is an entirely different experience.

OTOH, you have English and German and others that frequently use compound words, and the use of spaces becomes really important to understanding the writing.

I have zero experience with Thai.

link

hibbelig 1137 days ago

> OTOH, you have English and German and others that frequently use compound words, and the use of spaces becomes really important to understanding the writing.

Schleifmaschinenverleih would like to have a word with you.

This is parsed as Schleif-Maschinen-Verleih. Verleih means a rental company. The middle one is machine, and the first one I find both sanding and whetting as translations, not sure which one it is. So you can rent sanding and/or whetting machines there.

There are cases where it's ambiguous but for the most part the lack of spaces in compound nouns in German is not an issue.

A somewhat infamous example is Rohrohrzucker which should be parsed as Roh-Rohr-Zucker (raw cane sugar), but Rohr-Ohr-Zucker is also possible (pipe ear sugar). It's pretty clear when it happens that you got the wrong parsing but it takes a while to figure out what the right parsing is :-)

As far as speaking is concerned: I guess the extra spaces in English don't necessarily translate to pauses, do they? Is plugin pronounced differently from plug-in due to the hyphen?

link

zdragnar 1137 days ago

To clarify, I'm not saying that compound words are difficult to parse due to a lack of spaces. I'm saying that without any spaces in a sentence at all, it's harder to differentiate between compound and non-compound words.

blueberry vs. blue berry

stand up vs. standup

online vs. on line (northeast US term for queueing)

cartwheel vs. cart wheel

Stick compound words in a sentence that doesn't have any spaces at all, and you either have to pause to grok context, or context won't even help you (blue berry vs blueberry). At least German capitalizes all of its nouns, which would certainly help.

Compare this with Chinese, Korean, Japanese or other similar languages that don't use spaces at all (except perhaps after punctuation).

link

dumbotron 1138 days ago

> As an English native speaker who learned Mandarin, I really didn't find the lack of spaces harmful to learning the language.

Definitely. The logograms and being in a completely different language family are the real hurdles.

link

eric-hu 1138 days ago

No, it's different between Chinese and Thai.

Lexing is very clear in Chinese. It's never the case that you look at a Chinese sentence and don't know where a character ends and another begins. Take this sentence in both languages: "good morning, how are you"

早安，你好吗

This sentence clearly has "spaces" and I'm pretty sure any person illiterate in Chinese could tell you there are 5 characters / words. Technically the third character is composed of 人 and 尔 but I don't know that anyone, even kids or beginners, would mistake those as _not_ going together.

สวัสดีตอนเช้าคุณเป็นอย่างไรบ้าง

In contrast, Thai is as you say: lexing and parsing bleed together. There are 7 words in this sentence, but you need to lex the 10 syllables and run them through your mental dictionary to recognize the possible words they could be. My Thai is very limited, but there are examples of sentences out there that actually have multiple valid readings with different semantic meanings, depending on how you group sounds together.

link

Anon1096 1138 days ago

早安 is made up of 2 characters but is a single word. If you fall into the trap of thinking 1 character = 1 word, you won't understand a thing. In this case you'd have thought it meant "early safe" instead of "good morning".

link

eric-hu 1138 days ago

Okay, you make a good point. Let's look back at the GGP's comment though:

> Same with Chinese language, thus lexing and parsing requires knowing many more words than in languages with spaces between words.

In English, can you get away with knowing the meaning of "good" and "morning" and not "good morning", and know that I'm greeting you instead of commenting on the quality of this morning?

link

likpok 1138 days ago

Good morning is a bad example because it has a colloquial meaning that is a least a little idiomatic. Most other words/phrases in English don’t have this effect, while many Chinese words are like 早安. 了解, for example, can’t even be pronounced without correctly parsing the word.

link

eric-hu 1137 days ago

Okay, I concede that I may have forgotten that Chinese has its exceptions too. 了解 is indeed a good example. There are plenty in English though. Even with context, sometimes I have to really pause and think whether to pronounce read as red or reed (I read it just fine, I read English just fine).

Where I've had pain specifically with Thai is that I can't even know where a syllable begins and ends until I read a few "syllables" together and decide whether some vowels go with the consonant in front or behind it, and whether some an -ar should be pronounced as an -aan.

link

adastra22 1138 days ago

Chinese has pretty regular rules about grouping characters into words though, as most compounds are 2-characters, or a 4-character idiomatic phrase. Even if I know only half the characters in a sentence, I can usually guess the word boundaries correctly. It's not 100% reliable, but good enough to avoid confusion.

link

hnfong 1138 days ago

I guess it really depends on "dialect". Try that with Cantonese :)

As mentioned in another comment, single syllable words are much more common in Cantonese, and word combinations are much more "free" in the sense that there are a lot more ambiguity as to what counts as a "word" and what is merely two single-character-words idiomatically used together. There are also cases where grammatical constructs (and also foul words) are inserted in between a two-character word/idiomatic combo, and sometimes the characters are reversed, to the extent that it used to be a meme: https://evchk.fandom.com/zh/wiki/Y%E5%B7%B2x

It's gotten to a point where, after thinking about it for a couple years, I've come to believe that segmentation on Cantonese is a fool's errand...

Of course, there's also classical Chinese where most of the time a character is a word.

link

deadfoxygrandpa 1137 days ago

i think you're on to something about cantonese, but it's also true of mandarin. segmentation of words in chinese in general seems inherently messier than segmentation in english. also look at stuff like abbreviations: is 北大 one word? is it an abbreviation for 北京大学 the same way Caltech is an abbreviation for california institute of technology? is it just two single character words, each of which is an abbreviation? i think its much less clear than english

link

hnfong 1137 days ago

Segmentation in Mandarin is easier due to tendency of the language to use 2+ characters for words. With a high quality wordlist you will go a long way.

The problem with proper nouns is that they don't end up in dictionaries, same with slang and other terms that for reasons don't end up in dictionaries.

The additional problem with Cantonese is that there's a larger class of words where the constituent characters can move around as if they were words themselves. Even for a native speaker with some experience in lexicography, it can be difficult to determine word boundaries as there are many cases where a word with characters X+Y can be interpreted as just word X and word Y with some idiomatic meaning. This issue is more pronounced in Cantonese because there are more single character words in active use.

I've actually done this before. My experience is that naive segmentation on Mandarin text with wordlist is probably 80+% accurate, while using the same algorithm in Cantonese text (with cantonese wordlist) will definitely end up "wtf".

link

adastra22 1137 days ago

The same problem exists in Japanese FWIW, whose speakers like to make the same sorts of abbreviations despite not having a bisyllabic meter like Mandarin does. Japanese is somewhat helped by having multiple orthographies, however.

link