How do you read UTF-8 characters from an infinite byte stream - C#

c#, stream

Solution

Rather than `Encoding.UTF8.GetChars` which is designed to convert complete buffers, get an instance of `Decoder` and repeatedly call its member method `GetChars` this will make use of the `Decoder`'s internal buffer to handle partial multi-byte sequences from the end of one call to the next.

Problem

Normally, to read characters from a byte stream you use a StreamReader. In this example I'm reading records delimited by '\r' from an infinite stream. ``` using(var reader = new StreamReader(stream, Encoding.UTF8)) { var messageBuilder = new StringBuilder(); var nextChar = 'x'; while (reader.Peek() >= 0) { nextChar = (char)reader.Read() messageBuilder.Append(nextChar); if (nextChar == '\r') { ProcessBuffer(messageBuilder.ToString()); messageBuilder.Clear(); } } } ``` The problem is that the StreamReader has a small internal buffer, so if the code waiting for an 'end of record' delimiter ('\r' in this case) it has to wait until the StreamReader's internal buffer is flushed (usually because more bytes have arrived). This alternative implementation works for single byte UTF-8 characters, but will fail on multibyte characters. ``` int byteAsInt = 0; var messageBuilder = new StringBuilder(); while ((byteAsInt = stream.ReadByte()) != -1) { var nextChar = Encoding.UTF8.GetChars(new[]{(byte) byteAsInt}); Console.Write(nextChar[0]); messageBuilder.Append(nextChar); if (nextChar[0] == '\r') { ProcessBuffer(messageBuilder.ToString()); messageBuilder.Clear(); } } ``` How can I modify this code so that it works with multi-byte characters?

Original source