Unified lexical and token frontend
Design document · Original source
RFCs record designs and changes. A proposal appearing here does not mean its feature is ready to use. Explore current language support
Status recorded in the original: Proposed
Bound original source · SHA-256111442d48abcb140c2d83ce20191d28ae3227dd8f09d274f8992313579c2ec19
- Status: proposed
- Revision: 7
- Date: 2026-08-08
- Feature flag: none; internal shadow execution only in Stage A
- Depends on: none
- Stability boundary: the public
parsercapability remains stable
1. Purpose
N/M currently has two source readers. source-cst.ts produces a lossless token document for formatting, while parser-source-preprocessing.ts and parser.ts recognize semantic statements through line splitting and regular expressions. Every new structural feature therefore adds another source-reading path and makes exact columns, nested expressions, and recovery harder to preserve.
This RFC makes the CST tokenizer the sole lexical source and incrementally replaces line-regex recognition with a token-driven recursive-descent frontend. It defines no new language syntax. A conforming replacement accepts and rejects the same programs and emits the same public parse result as the current stable frontend.
2. Non-goals
Revision 1 does not:
- add compile-time expressions, generic circuits, loops, arrays, structs, or any other Wave 0.5 syntax;
- change
NMProgram,NMDiagnostic, source-line, qubit-map, constant, or parameter contracts; - change formatter output or make formatter diagnostics parser diagnostics;
- switch the production parser during Stage A;
- remove the legacy parser before the rollback window has elapsed.
3. Normative architecture
The target frontend has three ordered layers:
- Lexical layer.
parseNMSourceCSTis the only tokenizer. Tokens retain their exact UTF-16 offsets, zero-based line/column ranges, text, trivia, and stable source order. Printing every token must reproduce the input exactly. - Structural layer. A bounded recursive-descent reader consumes tokens, not source regular expressions, and creates statement/block ranges. String and comment tokens are indivisible; their punctuation never terminates a statement or changes nesting depth.
- Semantic lowering. Token ranges are lowered to the existing
NMProgramandNMDiagnostic[]contracts. During the final Stage A checkpoint this layer must not call the legacy frontend or use it as a result oracle.
The public parseNMCode(source, options) signature stays unchanged. Shadow entrypoints are internal evidence surfaces and are not a second supported language API.
4. Trivia, ranges, and recovery
- Whitespace, newline, line-comment, and block-comment tokens remain trivia; they are available to formatting and tooling but do not change semantics.
- Statement locations continue to use the current one-based public diagnostic line convention. Token positions remain zero-based internally.
- Semicolons terminate only at the current structural depth. A newline may terminate a semicolon-optional legacy statement only where the current grammar already permits it.
- Parenthesis, bracket, and brace nesting is bounded by the existing source and parser limits. The structural reader must be iterative or enforce an explicit maximum before recursive descent can exhaust the JavaScript stack.
- Stage A may record CST recovery issues out of band, but it must not expose a new diagnostic code or reorder stable diagnostics. New user-visible recovery behavior requires a later RFC revision.
5. Delivery stages
5.1 Stage A0 — lexical shadow
The CST token reader replaces only executable-line discovery in a shadow entrypoint. Existing semantic preprocessing/lowering is deliberately shared. The structural token tree, lossless reconstruction, executable-line parity, and full NMParseResult parity are measured. A0 is useful migration evidence but does not satisfy Stage A because semantic lowering is not independent.
5.2 Stage A1 — structural shadow
All module, directive, declaration, block, and executable statement boundaries come from the token reader. Regex helpers may still interpret the bounded text inside a token range, but no helper may split raw source into statements or infer nesting from raw characters. Imports and generated/expanded library source must use the same lexical reader.
5.3 Stage A2 — independent semantic shadow
The shadow frontend independently produces the complete NMParseResult. It must not invoke parseNMCode, the legacy preprocessor, or legacy statement dispatch. Deep equality across the complete corpus is the release condition. Only A2 with zero differences satisfies Stage A.
An independent semantic shadow may be delivered in numbered, closed slices before full A2 acceptance. Such a slice must return an out-of-band unsupported result for syntax outside its declared grammar, must never call or fall back to the legacy parser, and must report eligible and unsupported corpus counts separately. A slice label such as A2.0 is migration evidence, not permission to skip unsupported corpus entries and not a claim that Stage A2 has completed.
5.4 Stage B — controlled switch
After reviewed A2 evidence, the token frontend becomes the default. CLI/server deployments retain NM_LEGACY_FRONTEND=1 for one released version. Browser and Worker artifacts receive an equivalent build-time compatibility selection; a user program cannot select its own frontend.
5.5 Stage C — debt removal
After the rollback window, legacy source splitting, preprocessing dispatch, and obsolete regex code are removed. Architecture budgets must decrease; they may not be raised to accommodate two permanent frontends.
6. Parity corpus
Every parity report identifies the exact candidate SHA, Node version, corpus counts, option matrix, and difference count. The mandatory corpus is:
- every entry in
nmExamples(66 at revision 1), using its declared experimental options; - every entry in
nmStdLibrary(40 at revision 1), both directly and through a representative exact import; - every N/M source artifact discoverable under
tests/fixtures; - invalid/recovery fixtures for comments, strings, Unicode identifiers, semicolon-optional lines, inline blocks, CRLF, and nested delimiters;
- deterministic generated valid and invalid programs using
unit/fuzz-support.ts, with seed and minimized failure artifacts.
The harness fails if either baseline inventory count drifts without an explicit fixture update. A missing or skipped corpus class is a failure, not a warning.
7. Equality contract
For the same source and NMParseOptions, Stage A compares the full value:
program;- ordered
diagnostics, including code, severity, message, hint, translation data, line, and column when present; sourceLines;qubitMap,constants,params, andnumQubits.
Equality is structural deep equality with no diagnostic sorting, path normalization, numeric tolerance, or omitted undefined-bearing fields. A difference is recorded with a bounded source hash and structural path; reports must not retain private workspace source.
8. Performance and budgets
The default single-frontend parser is measured against benchmarks/nm-parser-baseline.json. Stage B cannot regress the committed p95 limit or RSS budget. Shadow mode measures its own overhead separately because running two frontends is evidence collection, not production latency.
check-architecture-budgets.mjs adds explicit budgets for the token frontend and parity harness. At Stage C the parser.ts line budget must decrease from its pre-migration value; it must never be increased as a migration shortcut.
9. Diagnostics and failure policy
Stage A introduces no user-visible parser diagnostics. The evidence harness may use these tool-only codes:
| Code | Meaning |
|---|---|
NM-FRONTEND-PARITY-001 | Full parse result differs |
NM-FRONTEND-PARITY-002 | Lossless token reconstruction differs |
NM-FRONTEND-PARITY-003 | Mandatory corpus or option coverage is missing |
NM-FRONTEND-PARITY-004 | Token nesting/statement bounds exceed the harness contract |
A parity failure is fail-closed for migration: it blocks Stage B but never changes the result returned by the stable parser.
10. Export and tooling
No JSON IR, OpenQASM, framework-export, simulator, or hardware contract changes in this RFC. Formatter, LSP, DAP, Playground, CLI, MCP, and Worker continue to consume the same stable parser result until Stage B. After the switch they must all use the same default frontend; embedding a separate editor grammar as a semantic authority is forbidden.
11. Acceptance
Stage A0 acceptance:
- CST printing is byte-for-byte equal for the mandatory corpus;
- token executable-line discovery equals the legacy line reader;
- the shadow entrypoint equals the complete legacy
NMParseResult; - the default
parseNMCodepath is byte-identical and does not execute the shadow frontend; - parser, semantics, runtime, CST, formatter, and formatter-fuzz suites pass.
Stage A1 acceptance:
- the shadow entrypoint obtains statement and block boundaries from CST tokens, including inline
if/elseand match-case bodies; - strings and comments never contribute delimiter or nesting transitions;
- workspace imports, exact-version standard-library source, generic-oracle declarations, generated generic source, and feed-forward source regions use the same structural reader;
- compatibility compounds preserve unsupported same-line legacy shells rather than silently adding syntax;
- the mandatory full-result corpus reports zero differences, and internal boundary metadata does not escape through
NMParseResult.sourceLines.
Stage A2/Stage B acceptance additionally requires:
- independent lowering with zero corpus differences across the option matrix;
- deterministic fuzz evidence with no unresolved minimized counterexample;
- parser p95/RSS within the committed baseline;
- an exact-candidate Node/OS matrix and reviewer approval;
- a tested legacy rollback selection for one release.
Until all additional conditions pass, RFC status remains proposed, the stable frontend remains authoritative, and no product surface may claim that the unified token frontend has shipped.
12. Revision history
- Revision 19 (2026-08-08): adds the dormant Stage B frontend-selection and rollback contract without changing the production default. The platform, not an N/M program, supplies the default.
NM_LEGACY_FRONTEND=1and the equivalentNEXT_PUBLIC_NM_LEGACY_FRONTEND=1browser/Worker build input force legacy precedence; malformed override values fail closed to legacy, and unsupported token syntax never triggers an automatic legacy fallback. Contract tests exercise token selection and both rollback channels. Activation still requires retained exact-candidate matrix artifacts and reviewer approval. - Revision 18 (2026-08-08): records the
A2.15performance-evidence slice. Lexical, structural, and executable-line lowering now share one CST instead of allocating three copies. The benchmark measures legacy and token candidates in separate processes against the same committed p95/RSS budgets, emits candidate/worktree-tagged JSON, and uploads both reports from every existing Windows/Linux and Node 20/22 matrix lane. A non-promotional local Windows Node 22.23.1 diagnostic measured token p95 46.00 ms and RSS delta 57,409,536 bytes against 90 ms and 120,795,955-byte budgets. Retained exact-candidate matrix artifacts, reviewer approval, and tested rollback selection remain open, so Stage A2 and the production switch remain unapproved. - Revision 17 (2026-08-08): records the
A2.14recovery and deterministic fuzz slice. The independent frontend now matches legacy behavior for all eight mandatory Unicode/CRLF, inline-block, comment/string, block-comment, unterminated-string, unbalanced-brace, semicolon-optional, and nested-delimiter classes across 17 fixed option runs. The shared deterministic fuzz harness exercises 5,000 nightly valid/invalid sources across stable and all-experimental profiles, reports its base and suite seeds, and writes a minimized replay artifact on failure; all 10,000 option runs have zero complete-result differences. Performance, exact-candidate matrix, reviewer, and rollback evidence remain open, so this is not complete Stage A2 and does not authorize the production switch. - Revision 16 (2026-08-08): records the
A2.13static-corpus completion slice. Independent lowering now covers bounded train blocks, oversized-QReg diagnostics, unresolved workspace imports and calls, mainless module/circuit documents, the frozen inline compatibility form, and token-level Turkish keyword normalization with preserved raw source. It has zero complete-result differences on all 192 stdlib/exact-import/fixture option runs, all 132 frozen example option runs, 5 hand-written cases, and 128 deterministic generated valid programs. The separately mandated invalid/recovery and generated-invalid fuzz corpus, performance, exact-candidate matrix, review, and rollback evidence remain open, so this is not complete Stage A2 and does not authorize the production switch. - Revision 15 (2026-08-08): records the
A2.12frozen-example completion slice. Independent lowering now covers learning metadata, the supported compatibilityif/elseand observable compounds, local unitary circuit declarations, AST-leveladjointexpansion, equivalence/fidelity/resource assertions, and a closed non-executable hybrid-function preview grammar. It has zero complete-result differences on 5 hand-written cases, 128 deterministic generated cases, 132 option runs over all 66 frozen examples, and 154 eligible runs in the 192-run stdlib/exact-import/fixture matrix. General invalid/recovery and deterministic fuzz acceptance, performance, and rollback evidence remain open, so this is not complete Stage A2 and does not authorize the production switch. - Revision 14 (2026-08-08): records the
A2.11runtime-directive and register-slice import slice. Independent lowering now covers bounded noise, mitigation, target constraints, surface estimates, and exact/stable circuit calls over fixed register slices. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 120 eligible option runs over 60 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Learning directives, user functions, compatibility compounds, and recovery forms stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 13 (2026-08-08): records the
A2.10bounded-control expansion slice. Independent token lowering now coversuntilplus the frozenrepeat, compile-timefor, and supportedcontrolledgate blocks, with deterministic AST and source-line expansion. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 108 eligible option runs over 54 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. General recovery/control forms and the remaining directives stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 12 (2026-08-08): records the
A2.9extended-predicate slice. An independent token parser now lowers bounded&&,||,!, grouped, and comparison predicates while preserving NM-RFC-0011 negotiation diagnostics; it does not call the existing feed-forward predicate parser. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 100 eligible option runs over 50 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Remaining control, directives, and recovery forms stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 11 (2026-08-08): records the
A2.8optimization/data slice and directive-module extraction. Independent lowering now covers the bounded@optimizeand@datasetforms used by the frozen examples plusencode(sample_row(...)), without legacy preprocessing. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 92 eligible option runs over 46 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Other directives and remaining structured syntax stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 10 (2026-08-08): records the
A2.7analysis-statement slice. Independent lowering now covers parameter-gradient statements and namedobserveblocks containing probability or named/inline expectation metrics, including exact nested source-line fidelity. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 88 eligible option runs over 44 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Other directives, extended control, and recovery forms stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 9 (2026-08-08): records the
A2.6fixed-size vector-parameter slice. Independent lowering expands validAngle[N]declarations into deterministic scalar parameters and resolves indexed parameter references in expressions and public source lines without invoking legacy preprocessing. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 84 eligible option runs over 42 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Invalid/recovery vector forms and the remaining structured syntax stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 8 (2026-08-08): records the
A2.5named-observable slice and modular import/gate extraction. Independent lowering now reads direct and exact-version observable packages, canonicalizes imported source lines, and resolves weighted Pauli sums in named expectation bindings and assertions. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 80 eligible option runs over 40 of 66 frozen examples, and 150 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Vector parameters, metrics, extended predicates, directives, and the remaining structured syntax stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 7 (2026-08-08): records the
A2.4syndrome and stable-control slice. Independent token lowering now covers syndrome measurements, Bit expressions, simpleifblocks, branch-invariant negotiation, inline expectation and entanglement assertions, and exact nested block source lines. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 78 eligible option runs over 39 of 66 frozen examples, and 134 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Extended predicates, named observables, directives, and the remaining structured syntax stay out of band, so this is not complete Stage A2 or production-switch evidence. - Revision 6 (2026-08-08): records the
A2.3algorithm-expression slice. Imported circuit calls now substitute and evaluate angle arguments while retaining exact source metadata. Token lowering also covers inline Pauli expectations, stable probability assertions, and bounded sample blocks. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 58 eligible option runs over 29 of 66 frozen examples, and 134 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Named observables, structured control, and the other unsupported corpus classes remain explicit, so this revision is not complete Stage A2 or production-switch evidence. - Revision 5 (2026-08-08): records the
A2.2reusable-circuit slice. The independent frontend now reads direct circuit-library documents, resolves stable and exact-version circuit imports, and expands parameter-free imported circuit calls without legacy parser dispatch. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 34 eligible option runs over 17 of 66 frozen examples, and 134 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Parameterized circuit calls and the other unsupported corpus classes remain explicit, so this revision is not complete Stage A2 evidence and does not authorize the production switch. - Revision 4 (2026-08-08): records the modular
A2.1independent slice. The cursor, program metadata, numeric-expression, and statement lowerers are separately budgeted and forbidden from importing legacy parser dispatch. A2.1 adds stable/all-experimental empty-feature metadata parity,@seed,@challenge,@target, scalar constants and parameters, the complete parameterized gate family, reset/barrier, and target-profile diagnostics. It has zero differences on 5 hand-written cases, 128 deterministic generated cases, 20 eligible option runs over 10 of 66 frozen examples, and 6 eligible runs in the 192-run stdlib/exact-import/fixture matrix. Unsupported corpus classes remain explicit, so this revision is not complete Stage A2 evidence. - Revision 3 (2026-08-08): defines fail-closed numbered Stage A2 slices. The first
A2.0implementation independently lowers the closed module/main, fixed-size QReg, fixed-arity gate, scalar measurement, and return subset. It has zero full-result differences on 3 hand-written cases, 128 deterministic generated cases, and 7 of the 66 frozen examples. The remaining examples, standard library, fixtures, option matrix, and recovery grammar are outside this slice, so this revision is not complete Stage A2 evidence. - Revision 2 (2026-08-08): records the Stage A1 structural-shadow contract and its evidence boundary. Semantic preprocessing and lowering are still shared; this revision is not Stage A2 or production-switch evidence.
- Revision 1 (2026-08-07): opened the proposal and Stage A0 lexical shadow.