How can I substitute Unicode characters with ASCII in Perl?
perl, unicode, utf-8
Solution
For a generic solution, Text::Unidecode transliterate pretty much anything that's thrown at it into pure US-ASCII.
So in your case this would work:
perl -C -MText::Unidecode -n -i -e'print unidecode( $_)' unicode_text.txt
The -C is there to make sure the input is read as utf8
It converts this:
l'été est arrivé à peine après aôut
¿España es un paìs muy lindo?
some special chars: » « ® ¼ ¶ – – — Ṉ
Some greek letters: β ÷ Θ ¬ the α and ω (or is it Ω?)
hiragana? みせる です
Здравствуйте
السلام عليكم
into this:
l'ete est arrive a peine apres aout
?Espana es un pais muy lindo?
some special chars: >> << (r) 1/4 P - - -- N
Some greek letters: b / Th ! the a and o (or is it O?)
hiragana? miseru desu
Zdravstvuitie
lslm `lykm
The last one shows the limits of the module, which can't infer the vowels and get as-salaamu `alaykum from the original arabic. It's still pretty good I think
Problem
I can do it in vim like so: ``` :%s/\%u2013/-/g ``` How do I do the equivalent in Perl? I thought this would do it but it doesn't seem to be working: ``` perl -i -pe 's/\x{2013}/-/g' my.dat ```