std/stream/bytes
std/stream/src/bytes.trb
Bytes, and the two stages between bytes and text.
Everything that comes off a file, a socket or a pipe arrives in chunks whose borders nobody chose: a line, a JSON
value or even a single character can be split across two of them. So the stages here are resumable - they keep what
is incomplete and hand it on when the rest arrives - and that is also why they are Stages and not methods: the
very same value works on an Iterate<Bytes> in a test.
alias Bytes
type Bytes = List<UInt8>
A chunk of bytes. An alias, not a type of its own: it is a List<UInt8> and every list operation works on it. The
name exists because "bytes" is what the signatures mean, and List<UInt8> in twenty places reads like arithmetic.
A chunk belongs to whoever produced it. A source hands one out and does not keep it, so a reader may hold on to
it - there is no borrowed buffer that has to be copied before the next await, because a list is a value.
Related
Sink.addText- writing text into a byte sink.
type Utf8Error
type Utf8Error with Show, Error
Bytes that are not UTF-8, with the offset of the byte that broke it. A String is never anything else.
field offset
offset: Int
Counted in bytes from the start of what was decoded. What that is belongs to the stage that failed and is said
there: the chunk for textOf, the line for lines, the whole stream for decodedText.
fn show
fn show(): String
invalid UTF-8 at byte {offset} - what a caller prints or logs.
fn lines
fn lines(): Stage<Bytes, Result<String, Utf8Error>>
Bytes to lines, split at \n, with a \r before it dropped so that CRLF files read the same as LF files. The line
break itself is not part of the line, and a last line without one is still a line.
Fallible, so the items are Results; source.through(lines()).checked() turns them into the stream's own failure.
The same stage works on a plain Iterate<Bytes>, with no Source and no Task involved.
Examples
const chunks = ["one\ntw".bytes().toList(), "o\nthree".bytes().toList()]
for line in chunks.through(lines()) {
print line
}
Errors
Utf8Errorfor a line whose bytes are not UTF-8, with the offset counted from the start of that line. One broken line fails once: every line around it arrives as it is, because each one is decoded on its own.
fn decodedText
fn decodedText(): Stage<Bytes, Result<String, Utf8Error>>
Bytes to text, one chunk at a time and cut at character borders: what arrives is whatever was complete, and a character split across two chunks is held back until the rest is there.
Examples
const chunks: List<Bytes> = [[0xE2, 0x82], [0xAC]]
print chunks.through(decodedText()).toList()
Errors
Utf8Errorfor bytes that are not UTF-8, with the offset counted from the start of the stream: every byte the stage has been handed since it started counts, not only the bytes of the chunk the failure was found in. A chunk border is nobody's choice, so an offset inside a chunk would name a byte the caller cannot point at.- At the end of the stream an incomplete sequence is such a failure too, and its offset is the byte that sequence starts on.
Pitfalls
- A sequence that is still incomplete when the stream ends is reported as invalid UTF-8, not delivered as a partial character: there is no more input coming to complete it.
fn textOf
fn textOf(bytes: Bytes): Result<String, Utf8Error>
A whole chunk of bytes as text - the short form for everything that has all of its bytes already (a file that was read in one piece, a body below its limit). Bytes that are not UTF-8 are a failure, never a replacement character.
Examples
const bytes: Bytes = [104, 105]
print textOf(bytes)
Errors
Utf8Errorwith the offset counted from the start ofbytes: the byte that cannot be decoded, or the byte an incomplete sequence at the end starts on. Nothing is held back here - there is no next chunk to complete it with.
fn encodedText
fn encodedText(): Stage<String, Bytes>
Text to bytes. UTF-8 needs no decision here: a String already is UTF-8.
Examples
const words = ["hi", "there"]
print words.through(encodedText()).toList()
Related
Sink.addText- the same step for one text, written straight into a byte sink.
extend Sink<Bytes, Failure>
extend<Failure> Sink<Bytes, Failure>
Writing text into a sink of bytes. An extension of the instantiated trait, so it is there for every byte sink - a file, standard output, a socket, the input of a child process - and nowhere else.
fn addText
var fn addText(text: String): Task<Result<Void, Failure>>
Encodes text as UTF-8 and writes it - the text form of add.
fn addLine
var fn addLine(text: String = ""): Task<Result<Void, Failure>>
\n, not the line break of the platform: a stream is bytes on a wire, not a text file of an operating system.