Skip to main content
NM-RFC-0027

Distribution and statistical assertions

Design document · Original source

RFCs record designs and changes. A proposal appearing here does not mean its feature is ready to use. Explore current language support

Status recorded in the original: Proposed

Bound original source · SHA-256
cc7009e6f13ed05e3aa5bfc9ec8ff88c1e88acb0b09707d0e525ba04a675933a

  • Status: proposed
  • Revision: 2 (2026-08-12)
  • Target contract: 0.3.0-experimental
  • Feature flag: experimental.distributionAssertions=true
  • Owners: language, runtime, verification, tooling, product
  • Depends on: NM-RFC-0003 deterministic property assertions
  • Does not change: existing assert probability, stable test semantics, backend ceilings, or hardware evidence
  • Implementation gate: this document defines a proposal only; section 16 approval is required before implementation

1. Summary

N/M currently supports deterministic point assertions, including finite-shot probability assertions with a user-selected tolerance. It does not yet let a test distinguish these two materially different claims:

  1. an analytic simulator distribution is numerically close to a target; and
  2. a finite sample supplies enough evidence for a bounded statistical claim.

This RFC proposes explicit analytic and sampled modes. They use distinct acceptance rules and produce distinct evidence. A sampled result is never reported as an exact probability, and an analytic result never carries a confidence level.

nm
@seed(17);
@shots(4096);

test "bell distribution" {
  let q: QReg<2> = qreg[2];
  H(q[0]);
  CNOT(q[0], q[1]);

  assert distribution(q[0..1]) analytic ~= {
    "00": 0.5,
    "01": 0.0,
    "10": 0.0,
    "11": 0.5
  } tvd <= 1e-9;

  assert probability(q[0..1] in {"00", "11"}) sampled ~= 1.0
    abs_error <= 0.05 alpha 0.01 family "bell";
}

The proposal is a bounded reproducibility and testing feature. It is not a device-certification, calibration, tomography, causal, or quantum-advantage claim.

2. Frozen decision set

Revision 2 proposes the following decisions as one reviewable unit:

  1. analytic and sampled are mandatory source keywords; no mode is inferred.
  2. Analytic assertions consume exact simulator probabilities and have no seed, shot, confidence, or p-value semantics.
  3. Sampled assertions require an unsigned 32-bit seed, 1..4096 shots, an explicit alpha, an explicit absolute-error budget, and a named family.
  4. Revision 2 uses conservative Hoeffding certificates and one closed Bonferroni family correction. It does not expose a menu of statistical tests.
  5. All sampled assertions in one named family are planned before execution; the family size cannot change after observing results.
  6. Every sampled assertion in one named family declares the same canonical alpha; a conflicting family declaration is rejected before sampling.
  7. Analytic distribution comparison uses total variation distance. Sampled distribution comparison reports empirical TVD but passes only when its simultaneous Hoeffding upper bound fits the declared TVD budget.
  8. Every result records the requested claim, estimator, uncertainty rule, seed, shots, family size, adjusted alpha, and limitations.
  9. The feature remains disabled and absent from capability manifests until all section 16 reviews are approved.

3. Closed source forms

ebnf
distribution-assertion =
  "assert" "distribution" "(" register-view ")"
  assertion-mode "~=" distribution-literal
  "tvd" "<=" probability-literal statistical-clause? ";" ;

event-assertion =
  "assert" "probability" "(" register-view "in" outcome-set ")"
  assertion-mode "~=" probability-literal
  "abs_error" "<=" probability-literal statistical-clause? ";" ;

assertion-mode = "analytic" | "sampled" ;
statistical-clause =
  "alpha" probability-literal "family" string-literal ;

statistical-clause is required in sampled mode and forbidden in analytic mode. Existing NM-RFC-0003 syntax remains unchanged and does not silently acquire these semantics.

register-view is a non-empty, compile-time-sized, disjoint view containing at most five qubits. Its written order defines bit-string order. The first listed qubit is the least-significant bit, matching N/M measurement artifacts.

4. Target distribution and event rules

A distribution literal is a closed map from canonical zero-padded bit strings to finite decimal probabilities. It must:

  • contain every one of the 2^width outcomes exactly once;
  • use lexicographic key order in its canonical artifact;
  • contain values in the inclusive range 0..1;
  • sum to 1 within 1e-12; and
  • contain at most 32 outcomes.

An event outcome set is non-empty, contains no duplicates, and uses only bit strings of the selected view width. Its expected probability is in 0..1. Expressions, wildcards, regular expressions, callbacks, runtime-generated sets, NaN, Infinity, and signed zero spellings are rejected.

5. Analytic semantics

Analytic mode is available only on a noiseless local statevector execution whose selected view contains at most five qubits. The runtime derives the marginal probability distribution from the final state without sampling.

For an event:

text
pass iff abs(actual_probability - expected_probability)
        <= absolute_error + 1e-12

For a complete distribution:

text
TVD(actual, expected) = 0.5 * sum_i abs(actual_i - expected_i)
pass iff TVD(actual, expected) <= tvd_budget + 1e-12

The result uses claimKind: "analytic-numeric". The word "exact" is not used because floating-point statevector execution is numerical, even though no finite-shot estimator is involved.

Analytic mode rejects trajectory noise, density-matrix proposals, hardware jobs, partial or post-selected histograms, and any backend that cannot expose the complete marginal distribution under the same contract.

6. Sampled semantics

Sampled mode consumes one terminal histogram generated under NM-RFC-0003's case-local deterministic random stream. @seed(uint32) and @shots(1..4096) are mandatory. Reusing or resampling until a test passes is outside the contract.

Let a named family contain H scalar hypotheses. An event contributes one hypothesis. A distribution contributes one hypothesis per outcome. The compiler computes H from all enabled sampled assertions with the same family name before execution. Every occurrence in that family must produce the same canonical binary64 alpha value and the same 17-significant-digit artifact serialization. A mismatch rejects the complete family before the first random draw; source order does not select a winner. 1 <= H <= 128 and:

text
adjusted_alpha = family_alpha / H
radius = sqrt(ln(2 / adjusted_alpha) / (2 * shots))

family_alpha is that unique family value and 0.001 <= family_alpha <= 0.1. The adjustment is Bonferroni and is the only revision-2 correction. Every logarithm and square root uses IEEE-754 binary64; artifacts serialize 17 significant decimal digits.

For an event with empirical frequency p_hat:

text
certified_error_upper = abs(p_hat - expected) + radius
pass iff certified_error_upper <= absolute_error

For a distribution with K outcomes and empirical distribution p_hat:

text
empirical_tvd = 0.5 * sum_i abs(p_hat_i - expected_i)
certified_tvd_upper = empirical_tvd + 0.5 * K * radius
pass iff certified_tvd_upper <= tvd_budget

The certificate is intentionally conservative. It provides a closed, deterministic false-positive bound rather than maximizing statistical power. If the requested shots cannot establish the claim, the assertion fails with an insufficient-evidence diagnostic; it does not relax the error budget.

7. Family planning and optional stopping

  • Family names are local to one test or one expanded property case.
  • An empty name, more than 64 UTF-8 bytes, or more than 16 families fails.
  • All hypotheses and their order are frozen before the first sample.
  • One family stores exactly one canonical alpha; every assertion references it instead of carrying an independently effective correction value.
  • A result cannot add, remove, reorder, or rename hypotheses.
  • Early success, repeated execution, best-seed selection, optional stopping, and selecting a family after viewing the histogram are not supported.
  • A property expansion derives an independent seed per case as specified by NM-RFC-0003; each case has its own family plan and evidence.

8. Runtime carrier and artifact

The compiler produces a closed nm-statistical-assertion-plan@0.1 carrier containing source identity, selected backend, ordered measured view, expected values, analytic or sampled mode, error budget, seed/shots where applicable, family plan, canonical family alpha, adjusted alpha, numeric policy, and a checksum over the complete plan and gate carrier.

Execution produces nm-statistical-assertion-report@0.1. Its closed records include:

  • plan, source, implementation, and execution identities;
  • backend and claimKind;
  • analytic probabilities or histogram counts, never a substituted mixture;
  • expected values, empirical values, TVD where applicable, and pass/fail;
  • seed, shots, random-source version, family, H, alpha, adjusted alpha, and Hoeffding radius for sampled claims;
  • insufficientEvidence as a distinct failure reason; and
  • limitations stating local simulation, no calibration binding, no hardware fidelity, and no causal or certification claim.

Raw source is not embedded. Readers recompute derived values and integrity. Reports are capped at 131,072 UTF-8 bytes and 128 scalar hypotheses.

9. Diagnostics

Table 1
CodeMeaning
NM-STAT-001experimental negotiation is absent
NM-STAT-002malformed mode, view, event, distribution, or numeric literal
NM-STAT-003target distribution is incomplete, duplicated, or does not sum to one
NM-STAT-004analytic assertion used with an incompatible backend or execution
NM-STAT-005sampled assertion lacks valid seed, shots, alpha, family, or error budget, or conflicts with its family's alpha
NM-STAT-006family, outcome, hypothesis, byte, or qubit bound exceeded
NM-STAT-007runtime carrier, histogram, measurement map, or integrity mismatch
NM-STAT-008analytic assertion failed its declared error budget
NM-STAT-009sampled evidence is insufficient for the declared error budget
NM-STAT-010report reader or derived-value recomputation failed

10. Tooling and product surface

After approval, Core, CLI, Worker, LSP, VS Code, Playground saved/share contexts, Test Explorer, and Experiment Studio must negotiate the same flag and publish the same claim labels. Test Explorer must visibly separate:

  • analytic numeric comparison;
  • finite-shot statistical certificate; and
  • insufficient evidence from an ordinary runtime failure.

Charts must show counts and uncertainty text in addition to color. A tooltip or histogram alone is not the evidence artifact. No UI may shorten a sampled claim to "exact probability" or "verified hardware distribution".

11. Export and compatibility

Strict JSON IR may preserve the closed plan after its schema and reader exist. OpenQASM, Quantikz, and framework circuit exports do not carry assertion semantics and must report that omission. Stable source and NM-RFC-0003 behavior remain byte-compatible while the new flag is false.

12. Security and resource policy

  • Validate source/artifact byte limits and every count before allocation.
  • Reject sparse arrays, inherited/accessor fields, unknown versions, and non-finite numbers.
  • Use one histogram for all assertions over the same terminal measurement; assertions cannot trigger hidden extra sampling.
  • Do not fetch data, packages, calibration, or provider results while reading a plan.
  • Telemetry records bounded identities/counts, never raw source or full private distributions.

13. Non-goals

  • Adaptive sequential tests, optional stopping, Bayesian posteriors, p-value menus, user-supplied statistical code, tomography, shadow estimation, or arbitrary observables.
  • Density-matrix trace distance, fidelity confidence intervals, readout-error mitigation, zero-noise extrapolation, or error-corrected logical claims.
  • Hardware certification, device comparison, advantage claims, or publication significance decisions.
  • Raising statevector, shot, property-case, or result-size limits.

14. Rollout sequence

  1. Freeze flag-off parser behavior, grammar, numeric fixtures, and schemas.
  2. Implement analytic event and complete-distribution comparisons.
  3. Implement sampled family planning and Hoeffding certificates.
  4. Add strict plan/report readers and Core/CLI/Worker parity.
  5. Add LSP, VS Code, Test Explorer, Playground, and Experiment Studio labels.
  6. Close cross-platform seeded conformance, hostile-carrier fuzz, browser, accessibility, and performance evidence.

15. Acceptance criteria

  • Analytic and sampled spellings cannot be confused or inferred.
  • The Bell fixture passes analytically at the frozen numeric tolerance.
  • Same seed and shots produce byte-identical sampled evidence on pinned Windows and Linux Node 22 runners.
  • Changing seed may change counts but never the family plan or uncertainty rule.
  • Two different canonical alpha values under one family name are rejected before histogram execution.
  • Removing a hypothesis after sampling invalidates the carrier.
  • A sampled assertion that lacks enough shots fails as insufficient evidence.
  • Unknown fields, oversized plans, malformed distributions, invalid counts, and forged derived values fail closed.
  • Every surface uses the same analytic/finite-shot language and limitations.

16. Required review record

Table 2
ReviewRequired decisionStatus
Language ownersource forms, view ordering, compatibilitypending
Runtime ownerhistogram reuse, seed stream, boundspending
Verification ownerTVD, one-alpha-per-family, and Hoeffding/Bonferroni semanticspending
Security ownerstrict carriers, allocation order, telemetrypending
Product ownerTest Explorer labels and claim boundariespending

A merge, prototype, passing test, or local UI is not approval. Until all rows are approved, the flag and both artifact names are proposed identifiers only. This revision creates no parser rule, runtime path, capability entry, or UI toggle.