std/text/lib
std/text/src/lib.trb
UTF-8 text and the Unicode scalar value it is made of: Char, one Unicode scalar value, and String, a sequence
of them stored as UTF-8. Bool belongs to std/core, not here - it is a primitive of the language, not a fact
about text.
type Char
native type Char with Equals, Compare, Hash, Show
A Unicode scalar value: one character, decoded from UTF-8.
Examples
const letter = 'A'
print letter.toLowerCase()
Related
String- text made of many of them.
fn isDigit
native fn isDigit(): Bool
Whether the character is a decimal digit: 0 to 9, and nothing else.
fn isLetter
native fn isLetter(): Bool
Whether the character is a letter. Exact for ASCII, and by the division Latin-1 itself makes above it.
fn isWhitespace
native fn isWhitespace(): Bool
Whether the character is whitespace: the ASCII ones, plus the White_Space code points a text really carries.
fn byteLength
native fn byteLength(): Int
The number of bytes of the character in UTF-8 (1 to 4): what it adds to an offset into a String.
fn toUpperCase
native fn toUpperCase(): Char
The same character in upper case, or itself where there is no single upper-case code point for it.
This is the simple case mapping: one code point in, one code point out, over ASCII, over the letters of
Latin-1 and over the one pair that reaches out of it (ÿ to Ÿ). So ß is answered unchanged - its upper case is
SS, which is two code points and does not fit a Char - and so is every code point the mapping does not cover.
Pitfalls
Upper-casing a text character by character is therefore not the same as a full Unicode mapping, and is what
String.toUpperCase does today.
fn toLowerCase
native fn toLowerCase(): Char
The same character in lower case, or itself where there is none. The mirror of Char.toUpperCase.
fn showNested
native fn showNested(): String
'A': inside of another value a character is written in single quotes, with escapes.
type TextIndex
type TextIndex with Equals, Compare, Hash, Show
A position in one text: before one of its characters, or at its end. Only a text hands one out - its searches, its
String.start and String.end, String.indexAfter - so a position inside of a character cannot be written, and
text[at..] with an at that text.indexOf answered cannot panic (docs/design/PANICS.md section 4). text[3..] is
a compile error: a number counted by eye is exactly what lands inside of a character.
It holds a byte offset, which is what makes slicing at one O(1) and lets the slice share the storage. There is no
arithmetic on it - at + 1 is the mistake the type exists to prevent, and String.indexAfter is what it means. A
number from outside - a file format, an editor - becomes one through String.indexAt, which checks it, and one goes
back to a number through String.byteOffset.
A position of one text used on another is a broken promise of the program, as an index of one list used on another is: slicing there panics where it is past the end or inside of a character of the text it is used on.
Examples
const text = "Grüße 👋"
match text.indexOf("ß") {
Some(at) => print text[at..]
None => print "no ß"
}
Related
String.indexOf- where most of them come from.String.dropping- a number of characters instead of a position.
type String
native type String with Equals, Compare, Hash, Show, Add, Slice<TextIndex>, From<Iterate<Char>>
UTF-8 text. There is no length() and no text[i] on purpose: say what you count (chars(), bytes()).
Positions come from searching and are a TextIndex: only the text hands them out, so slicing at one is O(1), shares
the storage and cannot land inside of a character. The byte-level members - byteLength, byteAt, sliceBytes,
byteOffsetOf - are for a format that counts bytes.
A String is always valid UTF-8. The only ways in are literals, slices at character boundaries,
String.from(Iterate<Char>) and runtime functions that validate - so reading a file whose bytes are not UTF-8 is
an IoError, and there is no replacement character anywhere in the language.
Examples
const greeting = "Hello, World!"
print greeting.toUpperCase()
print greeting.split(", ")
Related
fn from
static fn from(value: Iterate<Char>): String
String.from(characters), characters.to<String>().
Ordinary TorbScript, and it cannot be a function of the runtime for the same reason ArrayList.from cannot be
one: value is a trait-typed Iterate of the program, so reading it means calling iterate and next through
a witness table, which C code cannot do. So it collects the characters and hands them to concatenated: a String
is a value, so appending to one in a loop copies everything that is already in it, which is O(n²) in the bytes.
fn chars
fn chars(): Iterate<Char>
The characters, decoded one at a time: for character in text.chars(). Ordinary TorbScript, see Characters.
fn bytes
fn bytes(): Iterate<UInt8>
The bytes of the UTF-8, in order.
fn charAt
fn charAt(at: TextIndex): Char?
The character that begins at at, or None at the end. A position is never inside of a character.
fn charAtByte
native fn charAtByte(offset: Int): Char?
The character that begins at this byte offset, or None at and past the end: the byte-level form of String.charAt
for a format that counts bytes. An offset inside a character panics like every other bad offset.
fn start
fn start(): TextIndex
The position before the first character.
fn end
fn end(): TextIndex
The position after the last character.
fn indexAfter
fn indexAfter(at: TextIndex): TextIndex?
The position after the character that begins at at, or None at the end. O(1).
fn indexBefore
fn indexBefore(at: TextIndex): TextIndex?
The position before the character that ends at at, or None at the start. O(1): at most four bytes back.
fn indexAt
fn indexAt(byteOffset: Int): TextIndex?
The position at a byte offset, or None where the offset is past the end or inside of a character: how a number
from outside - a file format, an editor, a span of a lexer - becomes a position, checked.
fn byteOffset
fn byteOffset(of: TextIndex): Int
The byte offset of a position: for arithmetic on lines and columns, and for a format that counts bytes.
fn indexedChars
fn indexedChars(): Iterate<(index: TextIndex, character: Char)>
Every character with the position it begins at: what a scanner walks.
fn byteAt
native fn byteAt(offset: Int): UInt8?
The byte at this offset, or None at and past the end. A byte is never inside anything, so this never panics.
fn byteLength
native fn byteLength(): Int
How many bytes the text takes up in UTF-8. There is no shortcut for the number of characters: that means decoding.
fn isEmpty
native fn isEmpty(): Bool
Whether the text has no bytes at all. String.isBlank also treats pure whitespace as empty.
fn showNested
native fn showNested(): String
"a\nb": inside of another value a string is quoted, with escapes. show() is the text itself.
fn slice
fn slice(range: Bounds<TextIndex>): String
text[from..to] between two positions of this text. A position of another text that is past the end of this one
or inside one of its characters panics, and so does a start after the end.
What an open end means is the receiver's decision (std/core's Bounds says so), and for a text it is its
start and its end. An inclusive range (from..=to) takes the character that begins at to as well.
fn part
fn part(range: Bounds<TextIndex>): String?
The text the range covers, or None where it would panic - a position past the end or inside of a character,
which only a position of another text can be, or a start after the end: the total twin of text[from..to].
Examples
const text = "Grüße"
const at = text.indexOf("ß") ?? text.end()
print text.part(text.start()..at)
fn sliceBytes
native fn sliceBytes(from: Int, to: Int): String
The bytes from one offset to another. The one form the runtime has, and what slice decides its offsets for.
fn contains
native fn contains(part: String): Bool
Whether part occurs anywhere in the text.
fn startsWith
native fn startsWith(prefix: String): Bool
Whether the text begins with prefix.
fn endsWith
native fn endsWith(suffix: String): Bool
Whether the text ends with suffix.
fn indexOf
fn indexOf(part: String): TextIndex?
Where the first occurrence of part begins, or None if it does not occur. An empty part is at the start.
Examples
const path = "a/b/c"
if const Some(at) = path.indexOf("/") {
print path[..at]
}
fn lastIndexOf
fn lastIndexOf(part: String): TextIndex?
Where the last occurrence of part begins, or None if it does not occur. An empty part is at the end of the
text, which is where a search backwards finds it first.
Related
String.indexOf- the same search from the front.
fn byteOffsetOf
native fn byteOffsetOf(part: String): Int?
The byte offset of the first occurrence of part: String.indexOf for a format that counts bytes.
fn lastByteOffsetOf
native fn lastByteOffsetOf(part: String): Int?
The byte offset of the last occurrence of part: String.lastIndexOf for a format that counts bytes.
fn substringBefore
fn substringBefore(part: String): String?
Everything before the first occurrence of part, or None if it does not occur.
fn substringAfter
fn substringAfter(part: String): String?
Everything after the first occurrence of part, or None if it does not occur.
fn withoutPrefix
fn withoutPrefix(part: String): String?
The text without part at its start, or None where it does not start with it. It cannot panic: what it cuts off
is a whole text, so the cut is at a character boundary whatever the text holds.
Examples
print("--verbose".withoutPrefix("--") ?? "not an option")
Related
String.withoutSuffix- the same at the end.String.dropping- a number of characters instead of a text.
fn withoutSuffix
fn withoutSuffix(part: String): String?
The text without part at its end, or None where it does not end with it.
Examples
print("main.trb".withoutSuffix(".trb") ?? "main")
fn splitOnce
fn splitOnce(separator: String): (before: String, after: String)?
The text before and after the first occurrence of separator, which is dropped, or None where it does not
occur. One search, and no offset to get wrong - Go's strings.Cut.
Examples
match "name: Ada".splitOnce(": ") {
Some(parts) => print "{parts.before} is {parts.after}"
None => print "no field"
}
Related
String.split- at every occurrence.String.substringBefore,String.substringAfter- one of the two halves.
fn dropping
fn dropping(characters: Int): String
The text without its first characters characters. Total: a text with fewer characters answers the empty text,
and a count of zero or less the whole one, as skip of an Iterate does. O(n) in the count, not in the text.
A character is a Unicode scalar value, a Char of chars().
Examples
print "Grüße".dropping(characters: 3)
Related
String.prefix- the characters this drops.String.droppingLast- the same at the end.
fn droppingLast
fn droppingLast(characters: Int): String
The text without its last characters characters: the empty text where it has fewer, the whole one for a count of
zero or less. O(n) in the count.
Examples
print "Grüße".droppingLast(characters: 2)
fn prefix
fn prefix(characters: Int): String
The first characters characters of the text: the whole text where it has fewer, the empty one for a count of zero
or less, as take of an Iterate does. O(n) in the count.
Examples
print "Grüße".prefix(characters: 3)
Related
String.suffix- the last ones.
fn suffix
fn suffix(characters: Int): String
The last characters characters of the text: the whole text where it has fewer, the empty one for a count of zero
or less. O(n) in the count.
Examples
print "Grüße".suffix(characters: 2)
fn trim
native fn trim(): String
The text with leading and trailing whitespace removed.
fn toUpperCase
native fn toUpperCase(): String
The text in upper case: Char.toUpperCase for every character, and nothing more.
Every pair of that mapping needs as many bytes as the character it came from, so the result is exactly as long as
the text. A full Unicode mapping may make a text longer (ß becomes SS), and this one therefore leaves such a
character alone.
fn toLowerCase
native fn toLowerCase(): String
The text in lower case: Char.toLowerCase for every character. The mirror of String.toUpperCase.
fn replace
native fn replace(part: String, replacement: String): String
Every occurrence of part, replaced with replacement.
fn split
native fn split(separator: String): List<String>
The text cut apart at every occurrence of separator, which is dropped.
fn repeat
native fn repeat(times: Int): String
The text, written after itself times times: "ab".repeat(3) is "ababab".
Panics
When times is negative, and with a text longer than 4 GiB is not supported when the result would be longer
than a text can be.
fn isBlank
fn isBlank(): Bool
Whether the text is empty or holds nothing but whitespace.
fn lines
fn lines(): List<String>
The text cut apart at every \n.
extend Char with TryFrom<Int64, NumberRangeError>
extend Char with TryFrom<Int64, NumberRangeError>
Not every number is a Unicode scalar value (surrogates, everything above 0x10FFFF).
fn tryFrom
native static fn tryFrom(value: Int64): Result<Char, NumberRangeError>