Reference

std/stream/bytes

std/stream/src/bytes.trb

Bytes, and the two stages between bytes and text.

Everything that comes off a file, a socket or a pipe arrives in chunks whose borders nobody chose: a line, a JSON value or even a single character can be split across two of them. So the stages here are resumable - they keep what is incomplete and hand it on when the rest arrives - and that is also why they are Stages and not methods: the very same value works on an Iterate<Bytes> in a test.

alias Bytes

type Bytes = List<UInt8>

A chunk of bytes. An alias, not a type of its own: it is a List<UInt8> and every list operation works on it. The name exists because "bytes" is what the signatures mean, and List<UInt8> in twenty places reads like arithmetic.

A chunk belongs to whoever produced it. A source hands one out and does not keep it, so a reader may hold on to it - there is no borrowed buffer that has to be copied before the next await, because a list is a value.

Related

type Utf8Error

type Utf8Error with Show, Error

Bytes that are not UTF-8, with the offset of the byte that broke it. A String is never anything else.

field offset

offset: Int

Counted in bytes from the start of what was decoded. What that is belongs to the stage that failed and is said there: the chunk for textOf, the line for lines, the whole stream for decodedText.

fn show

fn show(): String

invalid UTF-8 at byte {offset} - what a caller prints or logs.

fn lines

fn lines(): Stage<Bytes, Result<String, Utf8Error>>

Bytes to lines, split at \n, with a \r before it dropped so that CRLF files read the same as LF files. The line break itself is not part of the line, and a last line without one is still a line.

Fallible, so the items are Results; source.through(lines()).checked() turns them into the stream's own failure. The same stage works on a plain Iterate<Bytes>, with no Source and no Task involved.

Examples

const chunks = ["one\ntw".bytes().toList(), "o\nthree".bytes().toList()]
for line in chunks.through(lines()) {
  print line
}

Errors

  • Utf8Error for a line whose bytes are not UTF-8, with the offset counted from the start of that line. One broken line fails once: every line around it arrives as it is, because each one is decoded on its own.

fn decodedText

fn decodedText(): Stage<Bytes, Result<String, Utf8Error>>

Bytes to text, one chunk at a time and cut at character borders: what arrives is whatever was complete, and a character split across two chunks is held back until the rest is there.

Examples

const chunks: List<Bytes> = [[0xE2, 0x82], [0xAC]]
print chunks.through(decodedText()).toList()

Errors

  • Utf8Error for bytes that are not UTF-8, with the offset counted from the start of the stream: every byte the stage has been handed since it started counts, not only the bytes of the chunk the failure was found in. A chunk border is nobody's choice, so an offset inside a chunk would name a byte the caller cannot point at.
  • At the end of the stream an incomplete sequence is such a failure too, and its offset is the byte that sequence starts on.

Pitfalls

  • A sequence that is still incomplete when the stream ends is reported as invalid UTF-8, not delivered as a partial character: there is no more input coming to complete it.

fn textOf

fn textOf(bytes: Bytes): Result<String, Utf8Error>

A whole chunk of bytes as text - the short form for everything that has all of its bytes already (a file that was read in one piece, a body below its limit). Bytes that are not UTF-8 are a failure, never a replacement character.

Examples

const bytes: Bytes = [104, 105]
print textOf(bytes)

Errors

  • Utf8Error with the offset counted from the start of bytes: the byte that cannot be decoded, or the byte an incomplete sequence at the end starts on. Nothing is held back here - there is no next chunk to complete it with.

fn encodedText

fn encodedText(): Stage<String, Bytes>

Text to bytes. UTF-8 needs no decision here: a String already is UTF-8.

Examples

const words = ["hi", "there"]
print words.through(encodedText()).toList()

Related

  • Sink.addText - the same step for one text, written straight into a byte sink.

extend Sink<Bytes, Failure>

extend<Failure> Sink<Bytes, Failure>

Writing text into a sink of bytes. An extension of the instantiated trait, so it is there for every byte sink - a file, standard output, a socket, the input of a child process - and nowhere else.

fn addText

var fn addText(text: String): Task<Result<Void, Failure>>

Encodes text as UTF-8 and writes it - the text form of add.

fn addLine

var fn addLine(text: String = ""): Task<Result<Void, Failure>>

\n, not the line break of the platform: a stream is bytes on a wire, not a text file of an operating system.