Windows UTF-8 to zero-terminated UTF-16 conversion helper.


utf16

(source -- result) Converts UTF-8 source to a zero-terminated Nat16 array.
  • Accepted source forms are StringView, Text, and String.
  • The whole source is interpreted as UTF-8 and converted to UTF-16 code units.
  • Unicode code points outside the Basic Multilingual Plane are represented by ordinary UTF-16 surrogate pairs.
  • The returned array owns its storage and is independent of the source lifetime.
  • Windows: conversion uses the complete byte extent of the accepted Text, StringView, or String, including embedded NUL bytes. The returned array size is the converted UTF-16 code-unit count plus one additional 0n16 terminator; subtract one for the payload code-unit count, not a Unicode-character count. A supplementary code point occupies two UTF-16 code units.
  • A zero character in the source remains part of the payload, so the result can contain earlier zeros or multiple trailing zeros. A Windows callee that treats the buffer as a zero-terminated string will stop at the first zero.
  • Windows: keep the returned array alive while a callee uses its buffer. Pass result.data to a typed Nat16-reference parameter, or result.data storageAddress to a Natx address parameter. Do not use result storageAddress as the UTF-16 buffer address. Reallocation, ownership transfer, or destruction can invalidate a retained buffer address.
  • For example, retain result: "hello" utf16; while a callee uses result.data storageAddress.
  • Conversion executes at run time even for known Text.

Failure conditions


Examples

Windows UTF-16 example

"String"          use
"control.Int32"   use
"windows/unicode" use

{} Int32 {} [
  result: "hello" utf16;
  ("size=" result.size LF) printList
  0
] "main" exportFunction

Expected Output

size=6

See also