Skip to main content
Language v0.2.0 · Preview

Stable QML Training Studio contracts

Run cancellable training, bounded trainability diagnostics, and fair optimizer benchmarks with versioned lifecycle, numeric, timing, and evaluation evidence.

Classical and quantum model comparison

Compare four fixed model families on the same held-out data: a majority baseline, logistic regression, RBF kernel ridge and a variational quantum classifier. This experiment uses the selected dataset, not the circuit in the editor.

Train, validation and test rows are disjoint. Scaling is fitted on training rows only. Validation selects hyperparameters; the test split is used for final reporting.

Accuracy is the fraction of correct labels. Balanced accuracy averages recall for the two classes. Signed-score MSE compares every model’s score with the same −1/+1 target; lower is better, and it is not probability calibration.

Deterministic baselines repeat the same fit; repeated identical scores are not independent evidence of low variance. Every run and seed is retained in the JSON evidence. Step budgets apply to iterative models; majority and kernel-ridge fits do not use gradient steps.

Load the comparison on the Quantum AI page and select circles_2d, spiral_2d, blobs_3f or parity_3bit. The default budget is 16 steps and 1 repeat, with at most 60 steps and 3 repeats. Inspect each model’s train, validation and test metrics, confusion matrices, validation-selected settings and separate resource counts. A successful experiment’s JSON includes every repeat, candidate, prediction, data split and scaling record. Changing settings or source invalidates the old result; failures and cancellation never produce downloadable partial evidence.

These are tiny deterministic teaching datasets. Scores do not establish quantum advantage or generalization to real data. Resource counts describe different kinds of work and cannot be compared as equivalent operations.

Load model comparison

Training Studio, trainability, and versioned benchmark contracts

These browser-facing QML surfaces now share bounded inputs, structured diagnostics, deterministic evidence, explicit evaluation accounting, and fail-closed asynchronous lifecycle rules.

v0.1
Studio contract
v0.1
Trainability contract
v0.2
Benchmark contract

Request-scoped Training Studio

Each workload has one current request id. A newer run supersedes the older run, stale progress or settlement cannot mutate the interface, cancellation is terminal, and consumer callback failures reject and terminate the Worker.

100
Maximum steps
8
Maximum parameters
200000
Maximum source characters
const run = client.train({ code, options: { steps: 40, learningRate: 0.2 } }, {
  onProgress: ({ step, cost }) => render(step, cost)
});
run.cancel(); // train-cancel + terminal rejection

Checked trainability profile

Seeded parameter-shift samples publish finite gradient variance and exact successful-sample and circuit-evaluation counts. Evaluation exceptions and non-finite values fail closed.

Up to 14 qubits, 8-400 samples, and 4 layers.

Fair optimizer benchmark

Every optimizer receives the same parsed source, seed, step budget, learning rate, objective, and initial parameters. Contract 0.2 adds frozen supervised train/validation/test evidence, a non-timing reproducibility hash, and versioned JSON/CSV export.

5 optimizers, 100 steps, 5 repeats, and 25 workloads maximum.

Race-safe lifecycle

Only the current request may append progress, publish results, report failures, or clear the running state. Old Promise settlements are harmless.

Exact evaluation evidence

Studio results carry requested steps, accepted progress events, terminal state, and circuit evaluations. Trainability uses exactly qubits×samples×2 evaluations.

Defensive timing

Benchmark clocks must be finite, monotonic, and non-negative. Throwing clocks, trainers, warmups, or progress consumers return structured failures.

Core, Worker, and docs parity

Core, CLI, Worker, dashboard downloads, and clean package exports share the same benchmark 0.2 evidence. Experimental NM-RFC-0009 exact-GD checkpoints remain a separately versioned artifact.

Stable QML product diagnostics

NM-QML-STUDIO-001NM-QML-STUDIO-002NM-QML-STUDIO-003NM-QML-STUDIO-004NM-QML-STUDIO-005NM-QML-301NM-QML-TRAINABILITY-001NM-QML-TRAINABILITY-002NM-QML-TRAINABILITY-003NM-QML-BENCH-001NM-QML-BENCH-002NM-QML-BENCH-003NM-QML-BENCH-004

Trainability thresholds are heuristic signals, benchmark durations are environment-specific, and local simulator results do not establish hardware performance or quantum advantage.

Hostile inputs, NaN/Infinity, callback failure, cancellation, stale settlement, exact Worker evidence, seeded profiles, fair optimizer matrices, clean package use, and bounded performance are automated promotion evidence.