Grammar Reference

This page defines the grammar of the current MPL compiler; semantics live on the Nodes page.

Notation

This page uses W3C EBNF. Quoted strings are literal text; backslash has no special meaning in the notation. ::= defines a production, adjacent terms concatenate, and | separates alternatives. ? * + mean zero-or-one, zero-or-more, and one-or-more. #xN denotes a hexadecimal code point; [a-z] is an inclusive range and [^...] excludes the listed characters. Eof uses $ as an end-of-input marker, not a literal dollar sign. Productions operate on Unicode scalar values decoded from valid UTF-8. Names and numbers are scanned greedily using the delimiter rules below, not split into adjacent valid prefixes.

Program

Source is an Item sequence ending at Eof. Group pairs nest Item sequences. Label and Variable are separate parser nodes; binding and lookup semantics are defined on the Variable node page.

Source ::= Item* Eof
Eof ::= $
Item ::= Node | Separator
Node ::= Group | Text | Comment | Number | Label | Variable
        | NameReadMember | NameWriteMember | NameMember
        | NameRead | NameWrite | Name
Group ::= Tuple | Block | Dict
Tuple ::= "(" Item* ")"
Block ::= "[" Item* "]"
Dict ::= "{" Item* "}"
Label ::= LabelName ":"
Variable ::= ";"
NameReadMember ::= ".@" NameBody
NameWriteMember ::= ".!" NameBody
NameMember ::= "." NameBody
NameRead ::= "@" NameBody
NameWrite ::= "!" NameBody

Tokens

Name, NameRead, NameWrite, NameMember, NameReadMember, and NameWriteMember must end before a NameDelimiter. An unprefixed spelling followed immediately by : forms a Label; the colon is forbidden in NameBody. Number must end before a Delimiter. The delimiter is not consumed by the token. A comma ends an existing token but can begin a BareName. A tab is a NameChar, not a separator. Text has quoted and nested guillemet forms. Escape accepts the six named escapes and the valid UTF-8 byte patterns in Utf8Escape.

Name ::= BareName | "!" | "@" | "."
LabelName ::= BareName | "!" | "@" | "."
NameBody ::= NameChar+
BareName ::= NameLead NameChar*
           | "-"
           | "-" NonDigitNameChar NameChar*
           | "," NameChar*
NameChar ::= [^#xA#xD !#x22#x23(),.:;@#x5B#x5D{}«»]
NameLead ::= [^#xA#xD !#x22#x23(),.0-9:;@#x5B#x5D{}«»-]
NonDigitNameChar ::= [^#xA#xD !#x22#x23(),.0-9:;@#x5B#x5D{}«»]
Delimiter ::= Eof | LineFeed | CarriageReturn | Space | "#" | ")" | "," | ";" | "]" | "}"
NameDelimiter ::= Delimiter | "."
Separator ::= Space | LineFeed | CarriageReturn
Space ::= #x20
LineFeed ::= #xA
CarriageReturn ::= #xD
Comment ::= "#" [^#xA]*
Number ::= Integer | Real
Integer ::= UnsignedInteger (IntSuffix | NatSuffix)?
         | "-" UnsignedInteger IntSuffix?
UnsignedInteger ::= DecimalInteger | HexInteger
DecimalInteger ::= "0" | [1-9] Digit*
HexInteger ::= "0x" HexDigit+
IntSuffix ::= "i" ("8" | "16" | "32" | "64" | "x")
NatSuffix ::= "n" ("8" | "16" | "32" | "64" | "x")
Real ::= "-"? RealBody RealSuffix?
RealBody ::= DecimalInteger (Fraction Exponent? | Exponent)
Fraction ::= "." Digit+
Exponent ::= "e" ("+" | "-")? ("0" | [1-9] Digit*)
RealSuffix ::= "r" ("32" | "64")
Digit ::= [0-9]
HexDigit ::= [0-9A-F]
Text ::= QuotedText | GuillemetText
QuotedText ::= '"' QuotedPart* '"'
QuotedPart ::= QuotedChar | Escape
QuotedChar ::= [^#x22#x5C]
GuillemetText ::= "«" GuillemetPart* "»"
GuillemetPart ::= GuillemetChar | Escape | GuillemetText
GuillemetChar ::= [^#x5C#xAB#xBB]
Escape ::= #x5C (#x22 | #x5C | [nr] | "«" | "»" | Utf8Escape)
Utf8Escape ::= U0000_007F | U0080_07FF | U0800_0FFF
             | U1000_CFFF | UD000_D7FF | UE000_FFFF
             | U10000_3FFFF | U40000_FFFFF | U100000_10FFFF
U0000_007F ::= [0-7] HexDigit
U0080_07FF ::= ("C" [2-9A-F] | "D" HexDigit) U80_BF
U0800_0FFF ::= "E0" [AB] HexDigit U80_BF
U1000_CFFF ::= "E" [1-9A-C] U80_BF U80_BF
UD000_D7FF ::= "ED" [89] HexDigit U80_BF
UE000_FFFF ::= "E" [EF] U80_BF U80_BF
U10000_3FFFF ::= "F0" [9AB] HexDigit U80_BF U80_BF
U40000_FFFFF ::= "F" [123] U80_BF U80_BF U80_BF
U100000_10FFFF ::= "F4" "8" HexDigit U80_BF U80_BF
U80_BF ::= [89AB] HexDigit

Integer without a suffix is signed Int32. IntSuffix accepts i8, i16, i32, i64, and ix; NatSuffix accepts n8, n16, n32, n64, and nx. RealSuffix accepts r32 and r64. Real without a suffix is Real64. Numeric range checks are not expanded into productions. The numeric magnitude in Exponent is limited to 308 for either sign. Conversion can reject a lexically valid Real that is out of range for its width, such as 9e308 or 1e-46r32.

Strict and permissive parsing

Strict parsing is the default. A CarriageReturn outside Comment and Text, including inside a Group, is rejected with Line contains carriage return; -permissiveParsing accepts it. Carriage returns inside Comment and Text are consumed by those productions. LF is the line-ending character counted by the parser.

Strict parsing rejects a Space immediately before LineFeed outside Comment and Text, with Line ends with space; -permissiveParsing accepts it. A Space immediately before Eof is accepted in either mode. These are the only two rules relaxed by the option. A tab remains part of NameChar and is rejected when a Number encounters it.

See also