Reference

std/text/lib

std/text/src/lib.trb

UTF-8 text and the Unicode scalar value it is made of: Char, one Unicode scalar value, and String, a sequence of them stored as UTF-8. Bool belongs to std/core, not here - it is a primitive of the language, not a fact about text.

type Char

native type Char with Equals, Compare, Hash, Show

A Unicode scalar value: one character, decoded from UTF-8.

Examples

const letter = 'A'
print letter.toLowerCase()

Related

  • String - text made of many of them.

fn isDigit

native fn isDigit(): Bool

Whether the character is a decimal digit: 0 to 9, and nothing else.

fn isLetter

native fn isLetter(): Bool

Whether the character is a letter. Exact for ASCII, and by the division Latin-1 itself makes above it.

fn isWhitespace

native fn isWhitespace(): Bool

Whether the character is whitespace: the ASCII ones, plus the White_Space code points a text really carries.

fn byteLength

native fn byteLength(): Int

The number of bytes of the character in UTF-8 (1 to 4): what it adds to an offset into a String.

fn toUpperCase

native fn toUpperCase(): Char

The same character in upper case, or itself where there is no single upper-case code point for it.

This is the simple case mapping: one code point in, one code point out, over ASCII, over the letters of Latin-1 and over the one pair that reaches out of it (ÿ to Ÿ). So ß is answered unchanged - its upper case is SS, which is two code points and does not fit a Char - and so is every code point the mapping does not cover.

Pitfalls

Upper-casing a text character by character is therefore not the same as a full Unicode mapping, and is what String.toUpperCase does today.

fn toLowerCase

native fn toLowerCase(): Char

The same character in lower case, or itself where there is none. The mirror of Char.toUpperCase.

fn showNested

native fn showNested(): String

'A': inside of another value a character is written in single quotes, with escapes.

type TextIndex

type TextIndex with Equals, Compare, Hash, Show

A position in one text: before one of its characters, or at its end. Only a text hands one out - its searches, its String.start and String.end, String.indexAfter - so a position inside of a character cannot be written, and text[at..] with an at that text.indexOf answered cannot panic (docs/design/PANICS.md section 4). text[3..] is a compile error: a number counted by eye is exactly what lands inside of a character.

It holds a byte offset, which is what makes slicing at one O(1) and lets the slice share the storage. There is no arithmetic on it - at + 1 is the mistake the type exists to prevent, and String.indexAfter is what it means. A number from outside - a file format, an editor - becomes one through String.indexAt, which checks it, and one goes back to a number through String.byteOffset.

A position of one text used on another is a broken promise of the program, as an index of one list used on another is: slicing there panics where it is past the end or inside of a character of the text it is used on.

Examples

const text = "Grüße 👋"
match text.indexOf("ß") {
  Some(at) => print text[at..]
  None => print "no ß"
}

Related

fn compare

fn compare(other: TextIndex): Ordering

The order of two positions in one text: the earlier one first.

fn show

fn show(): String

The byte offset, which is what a position is and what a message about one names.

type String

native type String with Equals, Compare, Hash, Show, Add, Slice<TextIndex>, From<Iterate<Char>>

UTF-8 text. There is no length() and no text[i] on purpose: say what you count (chars(), bytes()). Positions come from searching and are a TextIndex: only the text hands them out, so slicing at one is O(1), shares the storage and cannot land inside of a character. The byte-level members - byteLength, byteAt, sliceBytes, byteOffsetOf - are for a format that counts bytes.

A String is always valid UTF-8. The only ways in are literals, slices at character boundaries, String.from(Iterate<Char>) and runtime functions that validate - so reading a file whose bytes are not UTF-8 is an IoError, and there is no replacement character anywhere in the language.

Examples

const greeting = "Hello, World!"
print greeting.toUpperCase()
print greeting.split(", ")

Related

  • Char - one Unicode scalar value of the text.
  • Show - show() is the text unquoted, showNested() is quoted with escapes.

fn from

static fn from(value: Iterate<Char>): String

String.from(characters), characters.to<String>().

Ordinary TorbScript, and it cannot be a function of the runtime for the same reason ArrayList.from cannot be one: value is a trait-typed Iterate of the program, so reading it means calling iterate and next through a witness table, which C code cannot do. So it collects the characters and hands them to concatenated: a String is a value, so appending to one in a loop copies everything that is already in it, which is O(n²) in the bytes.

fn chars

fn chars(): Iterate<Char>

The characters, decoded one at a time: for character in text.chars(). Ordinary TorbScript, see Characters.

fn bytes

fn bytes(): Iterate<UInt8>

The bytes of the UTF-8, in order.

fn charAt

fn charAt(at: TextIndex): Char?

The character that begins at at, or None at the end. A position is never inside of a character.

fn charAtByte

native fn charAtByte(offset: Int): Char?

The character that begins at this byte offset, or None at and past the end: the byte-level form of String.charAt for a format that counts bytes. An offset inside a character panics like every other bad offset.

fn start

fn start(): TextIndex

The position before the first character.

fn end

fn end(): TextIndex

The position after the last character.

fn indexAfter

fn indexAfter(at: TextIndex): TextIndex?

The position after the character that begins at at, or None at the end. O(1).

fn indexBefore

fn indexBefore(at: TextIndex): TextIndex?

The position before the character that ends at at, or None at the start. O(1): at most four bytes back.

fn indexAt

fn indexAt(byteOffset: Int): TextIndex?

The position at a byte offset, or None where the offset is past the end or inside of a character: how a number from outside - a file format, an editor, a span of a lexer - becomes a position, checked.

fn byteOffset

fn byteOffset(of: TextIndex): Int

The byte offset of a position: for arithmetic on lines and columns, and for a format that counts bytes.

fn indexedChars

fn indexedChars(): Iterate<(index: TextIndex, character: Char)>

Every character with the position it begins at: what a scanner walks.

fn byteAt

native fn byteAt(offset: Int): UInt8?

The byte at this offset, or None at and past the end. A byte is never inside anything, so this never panics.

fn byteLength

native fn byteLength(): Int

How many bytes the text takes up in UTF-8. There is no shortcut for the number of characters: that means decoding.

fn isEmpty

native fn isEmpty(): Bool

Whether the text has no bytes at all. String.isBlank also treats pure whitespace as empty.

fn showNested

native fn showNested(): String

"a\nb": inside of another value a string is quoted, with escapes. show() is the text itself.

fn slice

fn slice(range: Bounds<TextIndex>): String

text[from..to] between two positions of this text. A position of another text that is past the end of this one or inside one of its characters panics, and so does a start after the end.

What an open end means is the receiver's decision (std/core's Bounds says so), and for a text it is its start and its end. An inclusive range (from..=to) takes the character that begins at to as well.

fn part

fn part(range: Bounds<TextIndex>): String?

The text the range covers, or None where it would panic - a position past the end or inside of a character, which only a position of another text can be, or a start after the end: the total twin of text[from..to].

Examples

const text = "Grüße"
const at = text.indexOf("ß") ?? text.end()
print text.part(text.start()..at)

fn sliceBytes

native fn sliceBytes(from: Int, to: Int): String

The bytes from one offset to another. The one form the runtime has, and what slice decides its offsets for.

fn contains

native fn contains(part: String): Bool

Whether part occurs anywhere in the text.

fn startsWith

native fn startsWith(prefix: String): Bool

Whether the text begins with prefix.

fn endsWith

native fn endsWith(suffix: String): Bool

Whether the text ends with suffix.

fn indexOf

fn indexOf(part: String): TextIndex?

Where the first occurrence of part begins, or None if it does not occur. An empty part is at the start.

Examples

const path = "a/b/c"
if const Some(at) = path.indexOf("/") {
  print path[..at]
}

fn lastIndexOf

fn lastIndexOf(part: String): TextIndex?

Where the last occurrence of part begins, or None if it does not occur. An empty part is at the end of the text, which is where a search backwards finds it first.

Related

fn byteOffsetOf

native fn byteOffsetOf(part: String): Int?

The byte offset of the first occurrence of part: String.indexOf for a format that counts bytes.

fn lastByteOffsetOf

native fn lastByteOffsetOf(part: String): Int?

The byte offset of the last occurrence of part: String.lastIndexOf for a format that counts bytes.

fn substringBefore

fn substringBefore(part: String): String?

Everything before the first occurrence of part, or None if it does not occur.

fn substringAfter

fn substringAfter(part: String): String?

Everything after the first occurrence of part, or None if it does not occur.

fn withoutPrefix

fn withoutPrefix(part: String): String?

The text without part at its start, or None where it does not start with it. It cannot panic: what it cuts off is a whole text, so the cut is at a character boundary whatever the text holds.

Examples

print("--verbose".withoutPrefix("--") ?? "not an option")

Related

fn withoutSuffix

fn withoutSuffix(part: String): String?

The text without part at its end, or None where it does not end with it.

Examples

print("main.trb".withoutSuffix(".trb") ?? "main")

fn splitOnce

fn splitOnce(separator: String): (before: String, after: String)?

The text before and after the first occurrence of separator, which is dropped, or None where it does not occur. One search, and no offset to get wrong - Go's strings.Cut.

Examples

match "name: Ada".splitOnce(": ") {
  Some(parts) => print "{parts.before} is {parts.after}"
  None => print "no field"
}

Related

fn dropping

fn dropping(characters: Int): String

The text without its first characters characters. Total: a text with fewer characters answers the empty text, and a count of zero or less the whole one, as skip of an Iterate does. O(n) in the count, not in the text.

A character is a Unicode scalar value, a Char of chars().

Examples

print "Grüße".dropping(characters: 3)

Related

fn droppingLast

fn droppingLast(characters: Int): String

The text without its last characters characters: the empty text where it has fewer, the whole one for a count of zero or less. O(n) in the count.

Examples

print "Grüße".droppingLast(characters: 2)

fn prefix

fn prefix(characters: Int): String

The first characters characters of the text: the whole text where it has fewer, the empty one for a count of zero or less, as take of an Iterate does. O(n) in the count.

Examples

print "Grüße".prefix(characters: 3)

Related

fn suffix

fn suffix(characters: Int): String

The last characters characters of the text: the whole text where it has fewer, the empty one for a count of zero or less. O(n) in the count.

Examples

print "Grüße".suffix(characters: 2)

fn trim

native fn trim(): String

The text with leading and trailing whitespace removed.

fn toUpperCase

native fn toUpperCase(): String

The text in upper case: Char.toUpperCase for every character, and nothing more.

Every pair of that mapping needs as many bytes as the character it came from, so the result is exactly as long as the text. A full Unicode mapping may make a text longer (ß becomes SS), and this one therefore leaves such a character alone.

fn toLowerCase

native fn toLowerCase(): String

The text in lower case: Char.toLowerCase for every character. The mirror of String.toUpperCase.

fn replace

native fn replace(part: String, replacement: String): String

Every occurrence of part, replaced with replacement.

fn split

native fn split(separator: String): List<String>

The text cut apart at every occurrence of separator, which is dropped.

fn repeat

native fn repeat(times: Int): String

The text, written after itself times times: "ab".repeat(3) is "ababab".

Panics

When times is negative, and with a text longer than 4 GiB is not supported when the result would be longer than a text can be.

fn isBlank

fn isBlank(): Bool

Whether the text is empty or holds nothing but whitespace.

fn lines

fn lines(): List<String>

The text cut apart at every \n.

extend Char with TryFrom<Int64, NumberRangeError>

extend Char with TryFrom<Int64, NumberRangeError>

Not every number is a Unicode scalar value (surrogates, everything above 0x10FFFF).

fn tryFrom

native static fn tryFrom(value: Int64): Result<Char, NumberRangeError>