Wide Char Yammering & My Great Buddy UTF-8
I'd been stuck a while on this, how to elegantly and correctly wrap text in lwl. I didn't use the curses library for it like most people would, so that I'd have to think through all the things I'm actually using, and how it all works.
I had assumed that everything would just be ascii-encoded and that I would worry about the problems with this assumption when I got to them. I pretty quickly Got To Them, as the internet uses UTF-8, not ascii, far and away as its dominant encoding. I had correctly identified that ascii would be a subset of whatever I would end up using: all "printing" characters in ascii are identical when encoded in UTF-8.
Outside of that, though, I found that UTF-8 diverges from ascii's one-column-per-character elegance by 1) introducing multibyte characters 2) who have variable widths when printed, of integers zero or more. Here are some examples of the kinds of things you will see in UTF-8 around the web:
| Hex | Description | Display | Width |
|---|---|---|---|
| 0x41 | Capital A | A | 1 |
| 0x031F | Combining Plus Sign Below | N/A | 0 |
| 0x01F360 | Yam Emoji | 🍠 | 2 |
| 0x306A | Hiranga Letter Na | な | 2 |
Given the diversity here, and the depth of the entire UTF-8 character encoding, it's definitely not super feasible or desireable to raw-dog this all, only using the single-byte-char IO functions that C provides.. So how exactly do you handle this mess in plain C?
As far as I can find, the cleanest and most widely compatible
way is to use the functions prescribed by wchar.h
from libc. It includes some very helpful functions, such as the following:
mbtowc: "multibyte to wide char", converts normal chars to wide chars
wctomb: "wide char to multibyte", converts wide chars to normal chars
wcwidth: get the (*approximate) number of columns a wide character advances the cursor
when printed in a terminal
getwchar: getchar but for wide characters
putwchar: putchar but for wide characters. Other stdio
functions also have "wide" alternatives like this too. Watch Out though: they don't play
nice with their standard variants! A fact which is not well documented in the relevant man pages...
Because of this, though, I prefer to use printf("%lc", wide_ch) instead of the devoted
wide IO functions.
wcwidth won't be exactly correct, especially
for emojis. There's not a great solution for this, other than using a better terminal
emulator like st.
Another sneaky gotcha, regarding wide characters in C, is the way all of this interacts with
locales. All of these functions' behavior changes based on your locale (you can use
echo $LANG to see what that is, if you want). More specifically the
LC_CTYPE category determines the exact encoding that C's wide-char methods
use. In our case, we want these functions to handle UTF-8. To ensure this, we can call
setlocale(LC_CTYPE, "") before all of our wide-char logic.
I wrote up some code to mess around with all this, which you can download here.
Ciao~~