UTF-16 code-unit and surrogate-pair helpers for reading and writing \uXXXX escapes. Import the module with "unicodeTools" use, or import one public name with "unicodeTools.NAME" use. This module supplies the UTF-16 escape operations used by Json; it complements windows/unicode, which converts UTF-8 to zero-terminated UTF-16, and String, which supplies UTF-8 text and string storage.


isHexDigit

(code -- valid) Tests one ASCII hexadecimal character code.

hexDigitToValue

(code -- value) Converts one hexadecimal character code to a value from 0 through 15.

readUnicodeCodeUnit

(reader -- codeUnit valid) Reads four hexadecimal character codes into one UTF-16 code unit.

catUnicodeCodeUnitHex

(result codeUnit --) Appends four lowercase hexadecimal digits to a String view.

isHighSurrogate

(codeUnit -- valid) Tests the inclusive high-surrogate interval 0xD800 .. 0xDBFF.

isLowSurrogate

(codeUnit -- valid) Tests the inclusive low-surrogate interval 0xDC00 .. 0xDFFF.

decodeSurrogatePair

(high low -- codePoint) Decodes a UTF-16 surrogate pair to a Unicode code point.

isSupplementaryCodePoint

(codePoint -- valid) Tests the inclusive supplementary-code-point interval 0x10000 .. 0x10FFFF.

encodeSurrogatePair

(codePoint -- high low) Encodes a supplementary code point as high then low surrogate values.

Unicode rules

Character codes, code units, code points, and surrogate values are Nat32, including values whose numeric range fits Nat16. Predicates return Cond; hexDigitToValue and decodeSurrogatePair return Nat32; encodeSurrogatePair returns high then low as Nat32 objects.

The returned codes, flags, and surrogate values are fresh computed stack objects, not views into the input. The surrogate and supplementary predicates can produce known results from known inputs.

Supply inputs in the documented ranges. Checked violations report invalid hexadecimal digit, high surrogate expected, low surrogate expected, or supplementary Unicode code point expected, as applicable. Known surrogate or code point violations can fail compilation; run-time assertions require DEBUG. Disabling checks does not make invalid input supported.

Hexadecimal character tests use the module's unknown ASCII-code values in this implementation, so even a known invalid character can fail only at run time.

catUnicodeCodeUnitHex is formatting, not validation: it masks the input to four hexadecimal digits. Convert a supplementary code point to its surrogate pair before formatting two UTF-16 escapes.


Examples

Hexadecimal code units

"String"       use
"control"      use
"unicodeTools" use

{} Int32 {} [
  result: String;
  "hex=" @result.cat
  @result 0x0041n32 catUnicodeCodeUnitHex
  " " @result.cat
  @result 0xD83Dn32 catUnicodeCodeUnitHex
  (@result.getStringView) printList
  0
] "main" exportFunction

Expected Output

hex=0041 d83d

Surrogate pair

"String"       use
"control"      use
"unicodeTools" use

{} Int32 {} [
  high: low: 0x1F600n32 encodeSurrogatePair;;
  ("high=" high " low=" low " decoded=" high low decodeSurrogatePair) printList
  0
] "main" exportFunction

Expected Output

high=55357 low=56832 decoded=128512

See also