Method
Every measured number on these pages comes from a circuit Qly ran on real hardware from its own account, through the same API anyone can use. This page says which circuits, what counts as a right answer, how a run is placed and recorded, and what the numbers cannot tell you.
1. The circuits
A Bell pair. Two qubits: a Hadamard, then a CNOT, then both measured. An ideal machine returns only 00 or 11. It is the editor's default example, which is why most machines have runs of it.
Mirror circuits (Proctor et al., arXiv:2112.09853). Random layers of one-qubit Clifford gates (H, S and their inverses) and CNOTs between neighbours, then an X on a random subset of qubits, then the same layers in reverse. An ideal machine returns one known bitstring. Ours are generated from a fixed seed (42) at widths 2 and 4 and depths 2, 8 and 32. Each answer below was checked two independent ways before any run: a statevector simulation of the exact circuit that was submitted, and Qiskit's stabilizer simulator. The first version of the generator had 10 of 12 answers wrong, and that check is what found it.
| Circuit | Right answer |
|---|---|
| 2 qubits, 2 layers | 00 |
| 2 qubits, 8 layers | 01 |
| 2 qubits, 32 layers | 00 |
| 4 qubits, 2 layers | 0001 |
| 4 qubits, 8 layers | 0000 |
| 4 qubits, 32 layers | 1001 |
Bitstrings are written with the highest-numbered qubit first, as Qiskit and Qly's results show them.
2. What a number means
A reading is the share of one run's shots that landed on a right answer: either Bell outcome, or the one mirror bitstring. That is all. It is not a fidelity. For the Bell pair in particular: measuring in the computational basis cannot see the relative phase, so this is an upper bound on how well the device made a Bell state, not a fidelity. A device producing an ordinary classical mixture of 00 and 11 — no entanglement at all — would score 100% here.
A run with fewer than 32 shots gets no reading at all, and the page says so rather than printing one shot as 100%. Runs are never averaged across days: a machine in June is not the same machine in September. Where a median appears, it is over the runs of one machine on one pair of qubits, and the shot count is printed beside every run so you can judge how much it rests on. There are no error bars on these pages; the shot count is the resolution.
3. Where a run lands, and what is recorded with it
A two-qubit circuit on a good pair and on a bad pair of the same chip are different experiments. Since 2026-09-22 every IBM run records the physical qubits it used, and campaign runs are pinned to named pairs and then checked against IBM's own record of the job, not just Qly's. Runs from before then, and on other vendors, say that the pair was not recorded. That is how IBM Kingston's low readings were traced to qubits 0 and 1: the same circuit on qubits 42 and 43, run minutes apart, read 98% to 99%.
Every run on real hardware since 2026-09-20 also carries a snapshot of the calibration the vendor had published when it was submitted. That is what lets a run be compared with the vendor's claim from the same moment, instead of with whatever the vendor publishes today.
Taking that apart one qubit at a time: on IBM Kingston, qubit 1 came back wrong on 38% to 48% of shots in all six mirror runs, including the shallowest, which is close to a coin toss. Qubits 2 and 3 of the same runs were wrong on under 6% at 2 and 8 layers. IBM published a 1.12% readout error for qubit 1 at the calibration all of these were submitted under. A repeat the next day, under IBM's following calibration, which put that figure at 2.49%, had qubit 1 wrong on 44.9% of shots. It fits the Bell pair on qubits 0 and 1 reading low elsewhere on this site; why qubit 1 behaves this way is not established.
A machine's median fidelity, as a vendor publishes it, is a claim about hundreds of pairs at once, and a two-qubit circuit uses one of them. On the Measured tab, some rows put what a vendor published for one exact pair beside the runs that landed on it, under that same calibration: on IBM Kingston, qubits 42 and 43 are the pair Qly's layout scoring rated best at that calibration, and 0 and 1 are where every earlier run landed. The runs alternated between the two, minutes apart, each under the calibration Qly recorded when it was submitted.
4. Whose runs these are
Only Qly's own. The measurements live in a fixed dataset in Qly's code, not in a query over the jobs people run, because a public page reading those would republish someone else's results. Where Qly's runs on a machine cannot be separated from another account's, the machine is listed as withheld, with the reason, rather than quietly left out.
5. What the vendors publish
Calibration figures are read live from each vendor's own API and labelled as the vendor's claims. They are medians, never means, because a few dead qubits drag a mean far from what a circuit sees, with the count of measured and unusable qubits and pairs shown underneath. The date shown is when the vendor says it characterised the machine, not when Qly asked.
6. What it costs
IBM bills machine time in whole seconds. Each of the 39 small jobs Qly ran on IBM from 2026-09-22 to 2026-09-24 was billed 2 seconds by IBM's own record. The 2 IBM campaigns behind these pages used 64 seconds in total (27 jobs on 2026-09-23, 54 s; 5 jobs on 2026-09-24, 10 s), read from IBM's usage counter before and after each.
7. How routing decides
Through Qly's API you can send a circuit with device: "auto" and one preference instead of naming a machine. Qly then picks the route and returns a receipt listing every machine it could have used: the number it compared for each, or the reason it left that machine out. It never scores machines or calls one best; it sorts, for that one request, by the single thing you asked for.
Price. The lowest charge for this circuit and shot count, computed the same way the charge itself is. For IBM machines, which bill by the second, that charge is an estimate, and the receipt says so.
Queue. The fewest jobs waiting ahead, as the provider reports it right now. A machine whose provider does not report a queue is left out rather than treated as empty; that is every IBM machine today.
Quality. The highest share of shots in the right answer, from Qly's own Bell-pair runs only, and only when they are current: at least 3 runs of 512 shots on the same pair of qubits, under the calibration the machine has now, from the last 2 days. A machine counts as behind the best only if the gap is at least 2 points and at least as wide as the best one's own runs disagree with each other; closer than that, they are treated as equal and the cheaper one is chosen. Today this preference always declines, and says why for each machine: no machine has current runs, and a standard-basis Bell count cannot see phase errors, so quality routing waits until Qly also measures in a second basis.
8. What these numbers cannot tell you
- Which machine is best. Nothing here is ranked, and a machine's result on one pair on one day does not describe the machine.
- Anything about quantum advantage. These are small circuits a laptop simulates instantly; they measure how often a machine returns a known answer, not whether it can do something a classical computer cannot.
- Anything about width eight or more, or about circuits unlike these two.
- Why a machine behaves as it does. A reading says what came back, not what caused it.
- The measured runs were not designed as an experiment. They are what people ran, which is why they are all the editor's default example, and why the shot counts and dates are uneven.
- Which physical qubits a circuit landed on was recorded for no run before 2026-09-22. Every IBM run before then that did not name a pair landed on qubits 0 and 1, because Qly places a circuit starting from qubit 0.
- Some machines Qly can reach have no measured run yet (see the Machines tab), and many runs here were taken before Qly recorded the calibration a run ran under.