Experimentally informed decoding of stabilizer codes based on syndrome correlations
Bibliographic record
Abstract
High-fidelity decoding of quantum error correction codes relies on an accurate experimental model of the physical errors occurring in the device. Because error probabilities can depend on the context of the applied operations, the error model is ideally calibrated using the same circuit as is used for the error correction experiment. Here, we present an experimental approach guided by an analytical formula to characterize the probability of independent errors using correlations in the syndrome data generated by executing the error correction circuit. Using the method on a distance-three surface code, we analyze error channels that flip an arbitrary number of syndrome elements, including Pauli <a:math xmlns:a="http://www.w3.org/1998/Math/MathML"> <a:mover accent="true"> <a:mi>Y</a:mi> <a:mo>̂</a:mo> </a:mover> </a:math> errors, hook errors, multiqubit errors, and leakage, in addition to standard Pauli <c:math xmlns:c="http://www.w3.org/1998/Math/MathML"> <c:mover accent="true"> <c:mi>X</c:mi> <c:mo>̂</c:mo> </c:mover> </c:math> and <e:math xmlns:e="http://www.w3.org/1998/Math/MathML"> <e:mover accent="true"> <e:mi>Z</e:mi> <e:mo>̂</e:mo> </e:mover> </e:math> errors. We use the method to find the optimal weights for a minimum-weight perfect matching decoder without relying on a theoretical error model. Additionally, we investigate whether improved knowledge of the Pauli <g:math xmlns:g="http://www.w3.org/1998/Math/MathML"> <g:mover accent="true"> <g:mi>Y</g:mi> <g:mo>̂</g:mo> </g:mover> </g:math> error channel, based on correlating the X- and Z-type error syndromes, can be exploited to enhance matching decoding. Furthermore, we find correlated errors that flip many syndrome elements over up to eight cycles, potentially caused by leakage of the data qubits out of the computational subspace. The presented method provides the tools for accurately calibrating a broad family of decoders, beyond the minimum-weight perfect matching decoder, without relying on prior knowledge of the error model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".