Skip to main content
Language v0.2.0 · Preview

Versioned N/M QML datasets

Inspect bounded embedded dataset schemas, permanent readers, disjoint train/validation/test splits, held-out validation evidence, and explicit provenance boundaries.

Embedded dataset 0.1 contract

Seven deterministic project fixtures share one strict schema, permanent reader, canonical serializer, curated split policy, and bounded supervised evaluation carrier.

v0.1
256 KiB
Pre-parse JSON limit
80
Maximum rows
8
Maximum features
Open the public dataset JSON Schema

Versioned embedded registry

Each registered dataset publishes exact feature names and immutable split counts. Runtime encoding and supervised training consume the same registry.

XOR 2D

xor_2d@0.1

Four angle-ready two-feature points for binary XOR-style QML encoding demos.

train: 2validation: 1test: 1

Moons 2D

moons_2d@0.1

Small curved two-feature sample for angle and reupload encoding examples.

train: 2validation: 1test: 1

Iris 2D

iris_2d@0.1

Tiny normalized sepal/petal slice for introductory feature encoding demos.

train: 2validation: 1test: 1

Circles 2D

circles_2d@0.1

Six inner-circle and six outer-ring points.

train: 6validation: 2test: 4

Spiral 2D

spiral_2d@0.1

Two deterministic eight-point spiral arms.

train: 10validation: 2test: 4

Blobs 3F

blobs_3f@0.1

Two compact clusters with three features.

train: 6validation: 2test: 4

Parity 3-bit

parity_3bit@0.1

All three-bit angle vectors labeled by parity.

train: 4validation: 2test: 2

Disjoint split invariant

Train, validation, and test are non-empty, disjoint, in range, and cover every row exactly once. Train retains both binary labels.

Held-out validation carrier

Optimization consumes only train rows. Final parameters are evaluated separately on validation rows and publish dataset version, split, sample counts, validation loss/accuracy, and actual circuit evaluations.

Permanent fail-closed reader

Malformed JSON, unknown root or nested fields, future versions, oversized payloads, and split tampering are rejected. Migration is explicit and lossless only.

Provenance and license boundary

These are project-generated deterministic teaching fixtures, not audited scientific or clinical datasets. Their metadata is descriptive and does not establish external benchmark validity.

Checked encoding boundary

Angle, reupload, amplitude, and IQP require a normalized finite state, unique in-range qubit indices, and 1–8 finite features. Invalid or non-finite output fails closed.

Stable dataset diagnostics

Dataset identity, shape, split, parse, unknown-field, limit, version, and feature-encoding failures use machine-readable codes independent of localized prose.

NM-QML-DATASET-001NM-QML-DATASET-002NM-QML-DATASET-003NM-QML-DATASET-004NM-QML-DATASET-005NM-QML-DATASET-006NM-QML-DATASET-007NM-QML-201NM-QML-202NM-QML-203NM-QML-204NM-QML-205

Held-out validation reduces training leakage but does not prove generalization. Test rows remain untouched by the training carrier; external data requires a separately reviewed schema, provenance, consent, and license policy.

All seven registries validate, the committed XOR fixture round-trips byte-for-byte, hostile JSON fails closed, split-aware encoding is tested, Core/CLI/Worker supervised results match, and the maximum 80×8 workload has a performance ratchet.