Are uppercase utf8 characters always the same number of bytes as their lowercase variants?
case-insensitive, unicode, utf-8
Solution
No.
Consider U+0069 "i" which has the octet value `69` in UTF-8. In the uppercase form U+0130 "İ" this code point forms the UTF-8 sequence `C4 B0`.
Obligatory note: case is locale-sensitive.
Problem
Obviously it is true for the latin alphabet. But I'm asking this in a conceptual sense, across languages and the Unicode spec. Practically this came up for comparing two strings. If you already know they aren't the same number of bytes—across all languages—can you consider that enough of a guarantee that they are not differently "cased" versions of the same string?