Hacker News new | ask | show | jobs
by avadodin 12 days ago
I have come to the conclusion that the only Unicode support needed in C is supporting pointers to char and arrays but lightweight C libraries are always welcome.
3 comments

(Author here) Dealing with encodings is already a big step when you aim to handle multiple ones and multiple OSes. You can see what I am talking about in the tests/ folder.

With Mojibake, I wanted to help people handle text by providing the smallest possible C/C++ library, without requiring them to use a +20MB library just to normalize a string, handle a flag emoji, or perform similar tasks.

See the CONFORMANCE_REQUIREMENTS.md file if you are interested in what the +17 versions of the Unicode standard have introduced.

Don't take my comment as dismissive of your project.

The people using it will probably have an easier time navigating Unicode text than they would have if they had used other existing libraries or tried to roll their own.

It's more a comment on the users who need to be warned that "no, you probably don't want your C program to know if that string actually fits in the 80 column terminal".

Keep in mind that I appreciated your comment. My library has a very narrow target

There is an MJB_FEATURE_CHARACTER_NAMES option you can set to zero if you don't want to have a function that returns the name of a codepoint, such as "LATIN SMALL LETTER E WITH ACUTE". This is something that probably most people do not need at all. This shrinks "Hello World" macOS ARM executable from 937KB to 663KB.

I should probably offer other runtime options so users can literally strip away everything they don't need. For example, as you suggested, measuring whether a string is less than 80 columns is something you don't do every day.

So how do you uppercase or lowercase an arbitrary character?

How do you check if it's a valid character?

How do you deal with combining characters?

Most of the world don't use just the Latin subset in ASCII or follow the same assumptions about letters, words, cases etc etc.

That's the trick.

You don't.

So you solution to dealing with X is 'dont deal with x'?
I guess you never have to deal with text if you think that’s enough? What kind of software do you write in C?
Software that treats Unicode(UTF-8) text as a bag of bytes, obviously. Even better if it is in a fixed length box of bytes.

A cursory look at the services provided by this library should dissuade people from attempting to work with Unicode text in C.

You can of course use or build a text processing engine's DSL that does all sorts of things people may want to do with Unicode text but C is hardly the best fit.