Formatting Guide
This page describes the source-code formatting conventions used across MPL sources. The rules cover visual layout, idiomatic expression choice, and the wording of compiler diagnostics; they are not language semantics. Most are conventions, with parser-sensitive whitespace cases stated below. mpl-format FILE reindents lines to two spaces per level, removes trailing whitespace and leading blank lines, collapses blank runs, normalizes the final line feed, removes blank lines after opening lines and before closing lines, inserts the space between a name and a following bracket and between a colon and a value, sorts and aligns consecutive use runs, and aligns a run of four or more binding rows whose values share one token shape, all in place; other column alignment is written by hand and kept.
Whitespace
- Indent with two spaces per level. No tabs: the parser does not treat a tab as whitespace.
- No trailing whitespace on lines. Outside comments and Text literals, the parser rejects spaces immediately before a line feed (see Grammar for
-permissiveParsing). - Separate logical sections inside a block with a single blank line.
- Put at least one space between a plain name and a following
(,[, or{;mpl-formatinserts one when none is present but preserves existing padding (rulespaceBeforeBracket). - End every non-empty source file with exactly one line feed.
mpl-formatadds a missing final line feed and removes extra trailing blank lines.
mpl-format FILE also removes leading blank lines and collapses runs of blank lines to one.
Scope closing
Closing a scope with ), ], or } ends one logical unit. The formatter removes a blank line immediately after an opening line or before a closing line; it never inserts one. Use a blank line to separate a closing line from a following non-closing logical step; consecutive closing lines need no blank.
Correct:
[
body
] loop
next step
Correct (consecutive closers, no blank required):
] if
] loop
] &&
Also accepted (the formatter preserves the missing blank line):
[
body
] loop
next step
Column alignment
When consecutive lines form a table of analogous items, the house convention aligns them into columns. The formatter rewrites each run of consecutive use lines itself (rule useStatementBlock): it sorts them by character code, pads the quoted names so use shares a column, and ensures one blank line follows the completed run when the next statement is not a closing line. At EOF or before a closing line it adds no blank. A run of four or more consecutive binding rows whose values share one token shape (four decimal numbers, four hexadecimal numbers, four Text literals, four Blocks of one shape) it aligns after the colons itself, token by token and inside a value too, adding or trimming padding: a decimal number on the last digit of its integer part, a hexadecimal number on its 0x (a sign extends left), every other token on its first character; a shorter run, or one whose values differ in shape, keeps its padding as written. In a binding row with a value it inserts one separator space after the colon when none follows it; a binding that takes its value from the stack stays name:;.
The following kinds of tables are aligned by convention:
- A sequence of
useimports, sorted by character code: pad each quoted name so the trailinguseis in one column. - A sequence of field or short-definition bindings with similar structure:
:stays attached to its name; pad after the colon so the right-hand sides start in one column. - A sequence of predicate-action pairs in a dispatch construct: align the predicate closing column, align any shared trailing keyword inside the action, and align the terminating operator of each pair.
- A sequence of struct or list items written on consecutive lines: align the comment column when each item has a trailing comment, and align the value column when each item has an explicit value.
Several snippets below are excerpts from mplc's parser rather than stand-alone programs. In those excerpts, names such as char, textValue, skip, fail, isDelimiter, AstNode, and processInteger are local or compiler-internal. The reusable helpers are ||, &&, between, and dup from "control" use; cond0, meetsAny, and contains from "algorithm" use; and assembleString from "String" use. Import those modules before using the helpers. In the short-circuit forms recommended below, || and && have stack effect (cond callable -- cond), with a callable that leaves one Cond. If the condition is unknown, both branches must compile with compatible outputs; if it is known, only the reached arm is compiled and no Cond requirement applies.
Aligned imports
"Array.Array" use
"Span.Span" use
"control.Cond" use
"control.between" use
Every quoted name is padded to the same width so use appears in one column.
Aligned bindings
isDecimal: ["0" "9" between];
isHex: [(@isDecimal ["A" "F" between]) meetsAny];
toDecimal: [.data 48n8 -];
toHex: [.data dup 65n8 < [48n8] [55n8] if -];
Each colon stays attached to its name; the padding after the colons aligns the opening [ and is kept as written, since these Blocks differ in shape.
Aligned dispatch descriptors
For predicate-action pairs in a dispatch construct, align three columns:
- The predicate's closing
]bracket. - Any shared terminal keyword inside the action (for example a final callable name at the end of each action).
- The shape of the action structure itself, so parallel elements across rows are in the same column.
(
[char "" = ] ["Unterminated escape sequence" fail]
[char "\"" = ] ["\"" @textValue.cat skip]
[char "\\" = ] ["\\" @textValue.cat skip]
[char "n" = ] ["\n" @textValue.cat skip]
[char "«" = ] ["«" @textValue.cat skip]
[char isHex ~] ["Invalid escape sequence" fail]
) cond0
When the action chain itself varies in length between rows (for example when some rows perform two checks and others perform four), pad the shorter chains so a shared downstream column (such as [isDelimiter ~] ||, the error arm, or the terminating if) aligns with the longer rows. Pad inside the descriptor brackets rather than after them.
Within any aligned block, never carry more than one column of spaces between two aligned columns. If several consecutive columns are spaces in every row, compress that run to a single separator column. Each row may then have its own internal run of spaces before the separator to reach the aligned column width.
(
[char "8" =] [skip ~ [isDelimiter ~] || ["short" fail] [0x7Fi8 AstNode.INT8 processInteger] if]
[char "1" =] [skip ~ [char "6" = ~] || [skip ~] || [isDelimiter ~] || ["short" fail] [0x7FFFi16 AstNode.INT16 processInteger] if]
[char "3" =] [skip ~ [char "2" = ~] || [skip ~] || [isDelimiter ~] || ["short" fail] [0x7FFFFFFF AstNode.INT32 processInteger] if]
) cond0
Between the hex constant and AstNode.INTn, only one column is all-spaces across every row; the varying space count on each side of that separator is the internal padding that each row needs to reach the aligned column's width.
Do not insert trailing spaces to pad a bare item inside brackets when there is no column to align against it. The default arm of a dispatch takes just its action with no padding.
Aligned comments on item lists
(
"INVALID" # An invalid class, used as a default value
"BLOCK" # A single code block
"COND" # An 8-bit conditional value
"INT8" # A generic 8-bit two's-complement integer
"INT16" # A generic 16-bit two's-complement integer
)
Each value is padded so the trailing # comment appears in one column. This applies whenever several consecutive lines each end in a comment about that line's item.
When alignment does not apply
- When only one line in a run has a given structure, padding to an imaginary column is not required.
- When two consecutive lines are not analogous (for example a dispatch descriptor next to a bare default action), each line stands alone.
Control-flow idioms
Express common two-branch patterns with short-circuit operators rather than full conditionals when possible.
| Intent | Prefer | Instead of |
|---|---|---|
Run body only if X is not TRUE. |
X [body] || |
X [TRUE] [body] if |
Run body only if X is TRUE and propagate a FALSE otherwise. |
X [body] && |
X [body] [FALSE] if |
Prefer the shorter form when the body is small and the intent is a plain short-circuit. Fall back to an explicit if when both arms carry substantial code.
Error-first convention
When a step tests a compound condition and chooses between an error/no-op branch and a work branch, place the shorter error/no-op branch first, attached to the condition. The condition is then phrased so that TRUE means "error".
Shape:
anyFailure [short-path] [work-path] if
When combining several sub-checks, build the failure condition with ~ on each sub-term and combine the terms with ||, so the chain evaluates to TRUE on any failure.
Avoid ending an &&-chain with a trailing ~. Replace that negated chain with ||-of-negations, keeping the if arms in the same order.
Incorrect:
A [B] && [C] && ~ [short-path] [work-path] if
Correct:
A ~ [B ~] || [C ~] || [short-path] [work-path] if
Multi-line expansion
When a branch carries substantial body, expand the dispatch onto multiple lines rather than packing everything onto one line. Nest each arm so the opening [ ends its line; a closing ] begins a line and is followed by the next arm opener or terminating token.
Incorrect (too compressed):
X [Y [char ":" =] || [short] [work with multiple statements] if] &&
Correct:
X [
Y [char ":" =] || [
short
] [
work
with multiple
statements
] if
] &&
Tail-token placement
Inside a multi-line block, put a trailing single-token loop/dispatch signal on its own line rather than appending it to a dense statement line. This makes the block's return value visible at a glance.
Incorrect:
[
statement
another
TRUE !flag FALSE
] ||
Correct:
[
statement
another
TRUE !flag
FALSE
] ||
Set-membership checks
When a value is compared against several alternatives, write a membership check against a small List rather than a chain of equalities combined with logical operators. This mirrors the intent and stays short when the set grows.
Incorrect:
char "a" = [char "b" =] || [char "c" =] ||
Correct:
("a" "b" "c") (char) contains
Error messages
This section sets the wording of the compiler's own diagnostics; new parser and semantic-compiler messages follow the shape below so they read uniformly regardless of the specific diagnostic.
Message shape
<who> <what>[[,] <extra>]
<who>is the phenomenon in question: a noun phrase naming the construct (for exampleText literal,Number,Real number,Line,Name,Member access, orModule) or a literal character quoted as«X».<what>is the verb phrase that describes the defect (for exampleis unterminated,overflows,starts with zero,has invalid suffix,is closed by «]»).<extra>is optional detail added to clarify the defect (for exampleexpected «"»,after «.»,in hex escape). Precede with a comma only when English grammar requires it — typically before a participial/clausal continuation likeexpected X. Plain prepositional phrases (in X,after X) take no comma.
Message examples
"Text literal is unterminated, expected hex digit"
"Text literal has non-hex character «G» in hex escape"
"Real number has exponent that starts with zero"
"Real number has non-digit character «a» after «.»"
"Number starts with zero"
"Number has hex digit «A» in decimal integer"
"Number overflows"
"Line ends with space"
"Line contains carriage return"
"«(» is closed by «]»"
"«(» is unterminated"
"Member access has no name after «@»"
Choosing <who>
- Prefer node-level phenomena. Say
Real number has exponent that starts with zero, notExponent starts with zero. - For errors where no enclosing node is meaningful (for example a stray token at top level), use the character itself as
<who>, quoted as«X».
Choosing <what>
- State the defect directly as a phrase describing the state. Avoid prohibition phrasing (
is not allowed,cannot have); the error context already implies wrongness. - Distinguish invalid and unexpected. An invalid character is one that is never valid in any context (for example a continuation byte in UTF-8 without a lead byte). An unexpected character is one that is valid elsewhere but not in the current position (for example
«[»inside a name).
Dynamic character inclusion
Include the offending character in the message when it speeds diagnosis. Use assembleString with « and » as visual quotes around the literal character:
("Name has unexpected character «" char "»") assembleString fail
For the » character itself — where «»» would be visually ambiguous — quote with ASCII double quotes:
"\"»\" has no matching opener"
Unterminated constructs
When a construct ends prematurely (for example EOF before an expected terminator), use the shape X is unterminated[, expected Y] where Y names the specific thing that was expected next:
"Text literal is unterminated, expected hex digit"
"Text literal is unterminated, expected escape character"
"Text literal is unterminated, expected «\"»"
"Text literal is unterminated, expected \"»\""
"«(» is unterminated"
For missing sub-parts inside a node, use has no X after «Y»:
"Member access has no name after «@»"
"Member access has no name after «!»"
Runnable examples
Formatted example
This runnable block applies the spacing and branch layout rules to a small conditional dispatch.
"String" use
"control" use
{} Int32 {} [
value: 3 dynamic;
value 2 > [
("large" LF) printList
] [
("small" LF) printList
] if
0
] "main" exportFunction
Expected Output
large