Quantum Error Mitigation: Useful Computation Before Fault Tolerance

Abstract technical illustration for the article: Quantum Error Mitigation: Useful Computation Before Fault Tolerance

Quantum computers are noisy. Gates misfire, qubits lose coherence, measurements misclassify. The accepted long-term fix is quantum error correction: encoding one logical qubit in many physical ones so errors can be detected and undone. But error correction is expensive — practical estimates run to a thousand or more physical qubits per logical qubit, well beyond what any current processor can spare. So here is the practical question: what can we compute with noisy hardware right now, before fault tolerance arrives?

The answer, increasingly, is error mitigation: a family of techniques that reduce the effect of noise on results without requiring error-corrected hardware. Mitigation does not make a computation fault-tolerant. What it does is give you a less biased estimate of the quantity you care about, at the cost of running more shots. The distinction matters, and the techniques are clever enough to be worth understanding on their own terms.

Mitigation is not correction

Error correction changes the computation: it spreads information across redundant qubits and actively repairs errors as they happen, so that in principle a computation of any length can succeed. Error mitigation changes the interpretation: you run the noisy circuit as-is, possibly several variants of it, and process the results to cancel out part of the noise's effect. The computation itself stays noisy; only the estimate improves.

The honest trade is this: mitigation buys accuracy with sampling. Every mitigation technique inflates the statistical variance of your estimator, so you need more repetitions — more shots — to reach a given precision, since sampling error falls as 1/√(shots). And mitigation does not scale to arbitrarily deep circuits: the sampling cost typically grows exponentially with circuit depth and noise level.

Zero-noise extrapolation: make it worse, then extrapolate

The most widely used technique is zero-noise extrapolation (ZNE), introduced by Temme and colleagues in 2017. The idea is delightfully indirect. You cannot run your circuit at zero noise, but you can run it at more noise — deliberately — and then extrapolate backwards.

Concretely: you evaluate the expectation value of your observable at the hardware's natural noise level, then again at amplified noise levels (typically noise factors of 1, 3, and 5). To amplify noise without changing what the circuit should compute, you use gate folding: replace a gate G with G·G†·G. Logically this is still G, but physically each gate adds its noise, so the folded circuit experiences roughly three times the noise. Then you fit the measured values and extrapolate to the zero-noise limit using Richardson extrapolation.

ZNE's great strength is that it needs no characterization of the noise — it is hardware-agnostic and simple to apply. Its weakness is that the extrapolation is only as good as your model of how the expectation value depends on noise. If the dependence is smooth and well-approximated by a low-order polynomial, ZNE works well; it is not guaranteed to be unbiased. The price is roughly a 3× shot overhead for the three noise levels, plus randomization repetitions. In Qiskit's runtime, ZNE is one of the standard resilience options: enabling it typically means running at noise factors [1, 3, 5] by default.

Probabilistic error cancellation: learn the noise, then invert it

Probabilistic error cancellation (PEC), also introduced by Temme and colleagues in 2017 and extended by Endo and co-workers in 2018, takes the opposite approach: instead of amplifying the noise, it cancels it. First you learn a model of the noise affecting each gate. Then you express the ideal (noiseless) gate as a quasi-probability distribution over noisy operations — a mathematical decomposition in which some coefficients can be negative — and sample circuits accordingly. The signed samples combine to recover an unbiased estimate of the ideal expectation value.

“Unbiased” is the key word: in principle, PEC can exactly recover the noiseless result. The catch is the sampling overhead. The quasi-probability decomposition carries a cost factor γ per gate, and the total overhead grows exponentially with circuit depth and total noise — roughly γ^(n·d) for n qubits and depth d. For shallow circuits with well-characterized noise, PEC is the most accurate mitigation technique available; for deep circuits at current error rates, the overhead becomes prohibitive. A 2023 Nature Physics paper by Kim and colleagues demonstrated that PEC produces competitive expectation values on noisy circuits, but the landmark 127-qubit experiment described below chose ZNE precisely because PEC's sampling overhead was too restrictive at that circuit volume.

Measurement error mitigation: the cheap win

Readout is often the noisiest step — measurements misclassify |0⟩ as |1⟩ and vice versa at rates of 1–5 percent. Measurement error mitigation (Bravyi et al., 2021) attacks this directly. You prepare each basis state, measure it many times, and build a calibration matrix M of “probability of measuring x given we prepared y.” Then you apply the inverse of that matrix to your experiment's counts to undo the readout distortion.

This is deterministic, fast, and nearly free — which is why it is enabled by default in most quantum runtimes. Qiskit's resilience level 1 applies TREX (twirled readout error extinction), a variant that twirls the readout noise into a simpler form before correcting it. If you are running on real hardware and not mitigating measurement error, you are leaving accuracy on the table for no reason.

The rest of the toolkit

Several more techniques fill out the mitigation landscape, each targeting a different error source:

  • Clifford data regression (Czarnik et al., 2021): replace the non-Clifford gates in your circuit with Clifford approximations to create training circuits that can be simulated classically (via the Gottesman–Knill theorem). A regression model learned on these maps noisy results toward ideal ones, and is then applied to the full circuit.
  • Virtual distillation (Huggins et al., 2021): uses the estimator Tr(ρ²O)/Tr(ρ²) — the observable's expectation over the “purified” state — which suppresses noise to second order in the error parameter. It needs only high shot counts, not extra circuits.
  • Dynamical decoupling (Viola, Knill & Lloyd, 1999): inserts π-pulse sequences into the idle windows of a circuit, where qubits would otherwise sit still accumulating dephasing noise. The pulses average out low-frequency noise. It is cheap, acts during scheduling rather than post-processing, and complements every technique above.

In practice these stack: a typical production run might use dynamical decoupling during transpilation, TREX for readout, and ZNE for gate error — Qiskit's resilience level 2 enables exactly this combination.

A real result: the 127-qubit Ising experiment

Mitigation is not just theory. In June 2023, Kim, Eddins, and colleagues at IBM Quantum, UC Berkeley, and Lawrence Berkeley National Laboratory published “Evidence for the utility of quantum computing before fault tolerance” in Nature (618:500–505). They simulated a kicked Ising model — a system of interacting spins — on the 127-qubit Eagle processor ibm_kyiv, which had median T1 and T2 times of 288 and 127 microseconds, unprecedented coherence at that scale.

The circuits were far beyond brute-force classical simulation: the full 127-qubit state vector cannot even be stored. Using zero-noise extrapolation, the team measured accurate expectation values for circuit volumes where leading classical approximations — matrix product states and isometric tensor network states — broke down. They established accuracy by comparing against exactly verifiable circuits at smaller scales. It was not a speedup for a useful application; it was something arguably more important: proof that a noisy processor, plus mitigation, can produce trustworthy answers in a regime classical methods cannot reach. Later work has continued the classical–quantum duel over this experiment, but the methodological point stands — mitigation turned a noisy 127-qubit device into a credible scientific instrument.

Using mitigation in practice

If you run circuits on real hardware through Qiskit, mitigation is one setting away. The runtime's EstimatorOptions expose a resilience level: level 0 disables mitigation, level 1 (the default) enables TREX readout mitigation, and level 2 adds ZNE on top. ZNE and PEC are mutually exclusive — one amplifies noise, the other cancels it — so you pick based on your circuit: ZNE as the general-purpose default, PEC when your circuits are shallow enough that its sampling overhead stays affordable.

Three practical rules follow from the theory. First, always mitigate measurement error — it is nearly free. Second, budget your shots for the mitigated estimator's variance, not the raw one; mitigation inflates error bars, and an under-shot mitigated estimate can look worse than the raw data. Third, know what mitigation cannot do: it will not rescue a circuit whose depth far exceeds what the hardware can support. The honest workflow is to shrink the circuit first — better transpilation, fewer two-qubit gates, dynamical decoupling in idle windows — and then let mitigation clean up what remains.

Fault tolerance is the destination, but it is years away. Error mitigation is what makes the journey productive: a principled way to trade the one resource we have in abundance — repetitions — for the one we lack — noiseless qubits. Every near-term quantum result you read about, from chemistry simulations to the 127-qubit Ising experiment, leans on these techniques. Understanding them is understanding how quantum computing actually gets used today.

Further reading

Similar Posts

Leave a Reply