Ask a measurable question
Does a small variational quantum classifier perform well on this dataset, under these settings, compared with simple classical alternatives? This is a useful educational question. It is much narrower than asking whether quantum machine learning is faster or better in general.
Quantum feature maps and quantum kernels offer ways to represent data and compare samples in a quantum feature space. The work of Schuld and Killoran explains this connection. It does not justify a blanket performance claim for every dataset or implementation.
N/M's Quantum AI workspace includes an educational comparison of a majority-class baseline, logistic regression, an RBF kernel ridge model and a two-qubit variational quantum model. This article describes how to read that comparison; it does not report a new benchmark result.
Separate fitting, selection and evaluation
Training data are used to fit parameters. Validation data guide the choice among candidate settings. Test data provide a held-out evaluation after that choice. Looking repeatedly at test scores while changing the model turns the test set into another tuning set.
The comparison runtime records the dataset and split hashes. Its preprocessing policy fits standardization on training rows, then applies the recorded transformation to the other splits. This is important: fitting a transformation on the whole dataset can leak information from held-out samples.
The implemented candidate-selection rule uses validation mean squared error, with a deterministic grid-order tie break. The training objectives can differ by model. A shared reporting metric does not mean that every model optimizes the same mathematical objective.
Run a comparison you can explain
- Open the comparison section and load the comparison tools.
- Select an educational dataset and keep the selected models fixed for the first run.
- Record the seed, repeats and training budget.
- Inspect the train, validation and test metrics separately.
- Save the result with its dataset, split and artifact hashes.
- Repeat with additional seeds before interpreting a small score difference.
A majority-class baseline is useful even when it looks trivial. If a complex model barely exceeds it, the experiment needs a different interpretation from one where a model consistently improves held-out prediction.
Read more than one score
Accuracy is the fraction of correct predicted labels. Balanced accuracy averages performance across the two classes when both are represented, making class imbalance easier to notice. The confusion matrix shows which errors occur. The reported signed-score mean squared error measures a score against labels encoded as -1 and 1; it is not a probability-calibration guarantee.
The artifact also records work counters, including kernel evaluations, statevector executions and related training operations. These counts describe different operations and should not be treated as directly interchangeable units of cost. The reproducibility hash excludes timing. A wall-clock measurement, if collected separately, needs the machine, browser, warm-up policy and workload recorded.
What this experiment can support
A held-out score can support a statement about that dataset, implementation and budget. It does not establish a hardware speedup: the quantum model here is simulated on classical hardware. Nor does a small synthetic classification fixture demonstrate success in drug discovery or an advantage on an industrial workload.
If the quantum model wins one run, ask whether the result persists across seeds, whether the baseline was reasonably tuned and whether the extra resources are justified. If it loses, the result is still useful: it exposes the limits of this model and provides a reproducible starting point for another hypothesis.
Sources
- Schuld and Killoran: Quantum Machine Learning in Feature Hilbert Spaces, Physical Review Letters, 2019.
- N/M Quantum AI comparison: the educational implementation discussed here.
- N/M documentation: language and local runtime scope.