Hacker News new | ask | show | jobs
by necovek 85 days ago
This is an article with a long introduction and then jumps straight to the point in one, final paragraph: Russia is abusing it for political messaging again. While yes, any tool will be abused like this, it really is also a tool to best codify spoken language of the Slavs (in a sense, it is trivially provable that Cyrillic script is better adapted even to languages which do not use it today, but have to resort to digraphs or glyphs with diacritics — some are thus not using it to distance from a particular influence instead).

None of the interesting bits of Cyrillic invention are covered, like how the original Slavic script was Glagolitic as the sibling mentioned, and only evolved into modern Cyrillic much later. Or how there was no lowercase until a few centuries ago, especially with the reform of Peter the Great.

With Slavic people, it's also worth noting that "Slav" actually means "word" or "letter" (of an alphabet), so legibility was part of the identity. In contrast, most Slavic people call Germans a variation of "Nemci", or mutes (those who cannot speak) — notably, most except Russians who call them Germans. Again, likely to distance themselves from the negative connotation with their aspiring historical partners.

7 comments

No idea where you're getting it from, Germans are Nemci in Russian as well. It's rather "unable to speak the language", meant for all foreigners but later stuck to Germans, presumably because German traders were the most common foreigners.
Apologies, it was mostly from running across different Russian maps with Германия that I took it as such (in Serbian it is Немачка). I stand corrected!

Nem/нем literally means "mute" in Serbian, perhaps it's a latter evolution per region either way.

It seems to me that you have entirely discredited yourself. You confidently make claims about the Russian language but don't even know the most basic thing about the point you were making.
It seems that I was wrong only partially, and this was totally not the core of my point (they do call the country Germaniya).

If making one mistake discredits the entire take, I'd hate to be your conversing partner ever.

It takes tremendous study to arrive at a conclusion and your inability to get _basic_ details correct, that your narrative depends on, betray that you have not done that study.

You are not my conversing partner. You are simply someone who emits nonsense onto the internet and people should be aware of you willingness to confabulate reality.

>"mute" in Serbian

Very far from Serbian only. Bulgarian, Russian, and even Balti-Slavic like Latvian is similar enough.

>Nem/нем literally means "mute" in Serbian,

Same in Russian

нем\немой - mute

немота - muteness

But yes, we do use Germany for country's name :)

Neamț (pl. nemți), pronounced something like neamtz/nemtzi is a German person in Romanian too.
> Germans are Nemci in Russian as well

I wanted to check; are you implying that Russian is not a Slavic language?

No, GP is saying that Russian uses the Latin root for Germans, I'm saying it doesn't. (it does for Germany though: "Germaniya").
I think I may have fallen victim to a GP midflight edit - I agree with you fwiw, it’s a stone cold fact.
"Slav" deriving from the Slavic term for "word" is something of a false etymology that was invented in the 19th century. It is implausible on philological grounds: you'd expect a different vowel in this word if this were the case, and the suffix *-ninъ is only otherwise used in terms derived from place names.

It is more likely[0] that the term derives from some toponym. This is in line with how tribal names tend to work in Europe and is not problematic in terms of historical linguistics, however it gives less fuel to romantic nationalism and armchair speculations about national "identities" or "mindsets".

-----

[0] https://en.wiktionary.org/wiki/Reconstruction:Proto-Slavic/s...

The irony for me being that when I was first learning Polish and looking for any and all mnemonics - “ah, that word is the number nine, and that one is ten because it has an s in the middle and that’s next to t for ten in the alphabet”-levels of desperate - the false etymology helped me set word, słowo, in my head, and the rather delightful dosłownie, literally / to the word, has remained ever since.

(tho while on the subject, it’s hard to beat wieloryb as a wonder that I don’t want to know the true etymology of ever because if there’s even a chance that the word for whale derived from the words great as-in-size + fish, I want to hang on to it forever)

Dunno. A nice parallel fact is that the word for "Germans" in at least a few Slavic languages literally means "mutes" - the ones who don't speak.

So you'd have the Slavs - the people of word - and the Germans - the mutes.

Exactly. In Polish "Niemcy" (Germans) comes straight from the mutes due to language barrier.
False etymology? You can roll back sound changes further to *ḱlew- in Proto-Indo-European

https://en.wiktionary.org/wiki/Reconstruction:Proto-Indo-Eur...

I always thought it was probable that it came from the same root as the word for "glory" - слава - as in, we're the glorious people.
> is trivially provable that Cyrillic script is better adapted even to languages which do not use it today, but have to resort to digraphs or glyphs with diacritics

Take a look at the Cyrillic section of Unicode to see your trivially provable claim being trivially disproven. You'll see all the same digraphs, glyphs, accents, graves etc. as used in Latin scripts.

It's also easy to see it easily disproven if you look at all the languages USSR forced cyrillic alphabet on.

To be fair, the parent post was clearly talking about Slavic languages, not "all the languages USSR forced cyrillic alphabet on", which were not Slavic and which required significant modifications to the alphabet.
Indeed: most notably, Croatian, Slovenian, Bosnian, Serbian and Montenegrin are all unambiguous with Cyrillic, but Latin script dominates, even in officially Cyrillic-first Serbia.

Again, it is seen as a political tool (pro-West or pro-Russia), when Cyrillic is technically better suited (there is certainly history as well, but that's very mixed up in the region).

Again, I am saying this as someone who has worked to implement things like full-text search, collation (lexical ordering/sorting) algorithms and tables, fonts and ligatures, functions like uppercase/titlecase/lowercase...

Eg. an already complex Unicode Collation Algorithm tables can never support exceptions with digraphs like "konjukcija" (nj is usually a digraph, but not here), etc.

The unique quirk with South Slavic languages is the linguistic work e.g. associated with Vuk Karadžić [1] which resulted in a cleaned up purely phonetic alphabet. This was done across the region and ended up getting plumbed through both alphabets, so e.g. the Croatians/Slovenes write in latin but with a handful of special characters for the unique sounds like "š" or the double-letter characters "dž" "lj", which also map 1-1 to stuff on the Serbian Cyrillic side.

It's the kind of legacy cleanup you love to see :-)

[1] https://en.wikipedia.org/wiki/Vuk_Karad%C5%BEi%C4%87

Serbia is still mostly Cyrillic though. It's a very interesting experiment since Croatia isn't and the languages are basically the same.
I invite you for a walk through Belgrade streets, maybe even with Google Street View. There will be Cyrillic in official signage, but ads and shop names will be predominately in Latin script. If there are some in Cyrillic, they are likely to be part of a newer "hipster" move to differentiate more for the tourists.
There are so many Russian émigrés in Belgrade that you hear Russian more than Serbian in the city center.
Since we're talking about Serbian below, here are some characters from Cyrillic Serbian Alphabet:

Ђ/ђ

Ћ/ћ

Љ/љ

Њ/њ

Џ/џ

Ј/ј

Various diacritical marks, digraph, a jod... What makes this Cyrillic more unambiguous than the Latin equivalents?

None of those are digraphs or have diacritics, each is a single letter/character

Compare with:

Š/S (Š = S + diacritic)

Nj (this "letter" is made up of two other letters)

> None of those are digraphs or have diacritics, each is a single letter/character

Okay, you got me, these don't have diacritics, but other Slavic languages do. Unicode committee decided that some of these are separate letters (Ѓ, Й, Ё etc.) and some are not (Ў), but still doesn't make these "trivially provable to suit Slavic languages better".

> Nj (this "letter" is made up of two other letters)

Indeed it is. Invented in 1818. E.g. Russian uses two "нь" for the same thing.

My point is, even if I may confuse my linguistic terminology from time to time, is... How does all this make Cyrillic "trivially provable" to be better suited for Slavic languages than Latin script? It's all the same: invent new letters or new letter combinations, or slap a few diacritics on top. And when that is not enough, borrow from Latin. E.g., j in Serbian, ї in Ukrainian and й in Russian for the same sound.

Some of this was a top-down overhaul of the writing system in the 19th century. Before that it was an awful mess and people just "vibe-wrote" the weird Slavic sounds using latin how they saw fit; try reading some old writings from that era. Or read some modern Polish or Czech text :-D
Most of the extra glyphs are for non-Slavic (Turk languages of Central Asia and Siberia). You see the same (and worse) in Latin Unicode pages — just look at how many variations of vowels 'a', 'i', or 'e' you have, consonants like 'c', 'z', 's'…
Even within Slavic languages there is plenty of weirdness: https://news.ycombinator.com/item?id=48064121
> it really is also a tool to best codify spoken language of the Slavs (in a sense, it is trivially provable that Cyrillic script is better adapted even to languages which do not use it today, but have to resort to digraphs or glyphs with diacritics — some are thus not using it to distance from a particular influence instead

I've heard this claim many times but never the reasoning behind it - by what metric is "ш" superior to "š" and so on?

It's less pronounced with diacritics, but enter Unicode normal forms: you can represent š either as š, or s followed by a diacritic. When you want to compare two strings, you have to normalize them to ensure you are comparing apples to apples. I can guarantee most software is broken in that regard. For Cyrillic, it just works.

With digraphs (lj, nj, dž + sometimes dj for đ too), it's even worse. Even capitalization is ambiguous: sometimes it's Lj and other times it's LJ. Then you have words like konjugacija where nj is not a digraph.

Interestingly — and not many know this — Unicode includes separate codepoints for all of the digraphs too. While well-intentioned, it only makes the problem worse.

Digraphs are especially sucky when you try sorting strings in a phonebook order as LJ comes after L, so you've got ...LI, LK..., LZ, LJA... With exceptions, it is even worse.

> It's less pronounced with diacritics, but enter Unicode normal forms: you can represent š either as š, or s followed by a diacritic. When you want to compare two strings, you have to normalize them to ensure you are comparing apples to apples. I can guarantee most software is broken in that regard. For Cyrillic, it just works.

It's the same with Unicode encoding of Cyrillic letters - й (U+0439) can be written as й (и U+0438 + ◌̆ U+0306)

> Interestingly — and not many know this — Unicode includes separate codepoints for all of the digraphs too. While well-intentioned, it only makes the problem worse.

Based on your description it seems that the root cause of the issues is using two letters to represent the digraph - for example N (U+004E) J (U+004A) instead of NJ (U+01CA) - and the sorting issues would be identical if people typed Н (U+041D) Ь (U+042C)instead of Њ (U+040A).

What's the reason for the digraph being substituted by 2 letters in the first case more often than in the second case?

You are absolutely right that there are examples where Cyrillic as used by Slavic languages is not perfectly "clean" either, and it's certainly a lot more nuanced than my simplistic and absolutist claim.

Perhaps people misunderstood me: it is not a _technical_ property of Cyrillic (vs Latin) script per se, but a combination of historical setting and ability to adapt the script to the (smaller) group's language. This has led to Cyrillic scripts being _developed_ to be technically more suitable for Slavic languages, because where Latin script was used, there was not as much liberty (perceived or real).

I mean, either is just a set of pictograms representing parts of spoken words, and obviously, if developed similarly, there is no difference between them. But for Slavic languages they were _not_ developed similarly, which is my point.

So, it's not "trivially provable that Cyrillic is better suited to Slavic languages". But that "the symbols representtion we settled on in software has some difficulties disambiguatuong some, but not all cases of symbol use in a language, a problem that is not unique to Slavic languages, see Dutch IJ, Turkish ı/i, German ß etc."
Decoupling choice of script from "symbols representation" is a weird approach — this is how people type them out.

Yes, problems are not unique to Slavic languages, but at least for _some_ Slavic languages, Cyrillic has been taken to the most simplified form that is _accidentally_ easy to process on a computer too.

But yes, I was a bit too absolutist, I agree — as ever, everything is more nuanced, so perhaps not "trivially provable", but in "closer to full differentiation in graphical representation while being simple and unambiguous to process on a computer"?

Most Latin-based scripts are just as unambiguous ;)
But the point was if this holds true for Slavic languages: my claim was it does not, as supported by the discussion.
Slav comes from slovo == слово which means word or speech, a.k.a slavs are people who can talk to each other which is a pattern in many other ethnic groups about differentiating between themselves and outsiders. Немци or mutes are those who cannot speak the language.
> Slavic people call Germans a variation of "Nemci", or mutes (those who cannot speak) — notably, most except Russians who call them Germans.

last time I checked we also call them "немцы" (Nemci and sounds exactly the same)

> some are thus not using it to distance from a particular influence instead

That's not the reason. The real reason is how those regions were Christianised - Cyril and Methodius created the first version of what would later evolve into cyrilic script and they were sent by Constantinople, while missionaries sent by Rome would use latin script.