No single number is enough.
The diagram shows a layered quantum-computer architecture and the relevant benchmarks associated with each part of the stack.
Orbsight industry briefing
A powerful tool for gaining clarity about the industry.
Today there are no quantum computers with a full stack capable of solving commercially useful problems at scale. But recently published technology and product roadmaps indicate there is good reason to believe that the first generation of quantum computers able to solve useful problems could emerge within the next 3 to 5 years1.
This makes today a sensible time to engage with quantum computing. The industry is maturing, value chains are forming and the roles of different companies in the industry ecosystem are becoming clearer. Uncertainty is still high, but opportunities abound. Some companies will find their niche in the value chain, others will develop their own vertically integrated solutions, and some will create new products and services using quantum computing.
This industrial growth highlights the need to separate the signal from the noise. Benchmarking is one of the best methods to do this. Used correctly, benchmarks can be used to measure progress, compare approaches, identify bottlenecks and assess whether quantum computing is moving towards genuine utility.
A chiselled horizontal line with a broad arrow marked underneath, used by the Ordnance Survey in the UK between 1840 and 1993. Benchmarks were made to provide a local reference to the height above mean sea level. The two benchmarks shown above highlight how one benchmark gives more information than the other.
No single benchmark can tell us whether a quantum computer is entirely good or bad, but different benchmarks provide insights. There are many benchmarks relating to quantum computing available in the literature2,3,4,5. A benchmark can relate to one or more elements of a quantum computer:
The product and higher layers tend to measure usefulness whereas the lower layers are more focused on capabilities. However, the relevance of each benchmark changes as the industry matures. Many benchmarks created during the early noisy intermediate-scale quantum (NISQ) era are not so useful now as we move into the fault-tolerant quantum computing (FTQC) era. Examples include GHZ and quantum volume. The best benchmarks will be well-defined, robust, efficient and technology-independent measures of performance. Here we describe some important benchmarks and discuss which ones are important today and which ones will become important as quantum computers approach utility and beyond.
The diagram shows a layered quantum-computer architecture and the relevant benchmarks associated with each part of the stack.
The end-user rarely cares what is inside the product. Product-level benchmarks therefore resemble those used in classical computing: energy per valuable solution, operational throughput, cost and physical footprint.
The applications and algorithms layers identify valuable problems and translate them into quantum computations that could outperform practical classical approaches. Useful benchmarks include time-to-solution, answer quality, cost per solution, comparison with classical baselines, problem size, qubits to solution, circuit depth and operation count. These measures connect technical progress to valuable outcomes.
The middleware contains the control and fault-tolerant functions needed to run reliable quantum computations. Compilation, fault-tolerant orchestration, logical operations, quantum error correction, control, readout and calibration form an interconnected middleware. These layers often share hardware and software resources, so their benchmarks are rarely isolated.
This layer contains the physical qubits and the hardware needed to support them. Relevant benchmarks include physical-qubit count, single- and two-qubit gate fidelity, coherence and leakage. Better physical qubits reduce the overhead and computational burden imposed on every layer above them.
Regardless of whether we treat the quantum computer as being a stand-alone unit or being integrated into a classical high-performance computer (HPC), we treat the computer as a ‘black box’ at the product level. It’s important to understand that the end-user rarely cares what’s inside the product. This means suitable benchmarks have more in common with classical computers.
Consider what makes the product useful. A quantum computer becomes useful when its end-to-end performance offers a meaningful advantage over the best practical classical alternative. From this understanding, the total energy per valuable solution must be considered an important benchmark. In the age of commercially useful quantum computers, this benchmark will come to the fore. Other meaningful product-level benchmarks include operational throughput, the cost and footprint of the product, including qubit density - qubit/m3 and qubit/kg.
There will be no value in quantum computing if no appropriate applications can be identified, and if no useful quantum algorithms can be devised. Fortunately, neither appears to be an insurmountable obstacle. Many major industries could benefit from computations that improve accuracy, speed, optimisation or simulation capability beyond what is practical with classical computers alone6. The prospect that quantum computers could solve problems that are intractable using conventional computing amplifies their disruptive potential. Several famous algorithms such as Shor’s factoring and Grover’s search algorithm indicate that quantum algorithms will have commercial and strategic value. The lack of a fully functional quantum computer certainly hinders the quest for all the possible useful algorithms, but the foundations for useful quantum algorithms are being laid7. DARPA’s Quantum Benchmark Initiative (QBI) is an independent verification and validation programme for utility-scale quantum-computer architectures. This work is also helping steer the development of suitable computing architectures to efficiently execute these algorithms8. Useful benchmarks are those that measure problem size, time-to-solution and comparison with classical baselines. The number of qubits (qubits to solution), circuit depth, and the number and types of operations needed to run each algorithm are all important values. There is already provision for obtaining this information from several sources including FTCircuitBench9 and BenchQC4. The Quantum LINPACK/RACBEM is based on the LINPACK benchmark used in supercomputing and is useful for benchmarking linear algebra and high-performance computing (HPC)-style workloads10. In addition to providing information about where quantum computers will be used in early commercial applications, these benchmarks will be valuable for investors and analysts.
Consider the evolution of published resource estimates for factoring RSA-2048 using Shor's algorithm11. This algorithm is one of the best-known examples of how a fault-tolerant quantum computer could break widely used public-key cryptographic systems. Using current classical algorithms and realistic computational resources, factoring RSA-2048 would probably take longer than the age of the universe. But a quantum computer using Shor’s algorithm could solve the same problem in hours or days, depending on its architecture and implementation. The two relevant benchmarks are the number of qubits needed (qubits to solution) and time to solution.
Early realistic surface-code resource estimates suggested that factoring RSA-2048 could require roughly one billion physical qubits and months or years of runtime. The estimate established an important baseline for fault-tolerant resource analysis.
Gidney and Ekerå dramatically reduced the estimate through improved modular arithmetic, magic-state factories and surface-code modelling. The result became one of the best-known reference points for Shor’s algorithm.
Further algorithmic and error-correction improvements reduced the estimated qubit requirement by more than twenty times, while allowing a longer runtime of less than one week.
A reconfigurable neutral-atom proposal used high-rate codes and different architectural assumptions to reach a much lower qubit estimate. The runtime is longer, so this point illustrates a space–time trade-off rather than a like-for-like improvement.
The number of qubits needed has fallen by several orders of magnitude, but time to solution has become an architectural trade-off. Optimising this balance will become a major driver in the design of future fault-tolerant quantum computers.
Published resource estimates
| Year | Authors | Estimated qubits to solution | Estimated time to solution |
|---|---|---|---|
| 2012 | Austin Fowler et al. | ~1 billion | Months–years |
| 2019 / 2021 | Craig Gidney & Martin Ekerå | 20 million | ≈8 hours |
| 2025 | Craig Gidney | <1 million | <1 week |
| 2026 | Cain et al. | ~10,000 (architecture-specific) | Longer than the surface-code estimates |
The middleware level runs the quantum algorithms, containing the control and fault-tolerant functions. Here, the middleware is displayed as a hierarchy of functions and benchmarks to help visualise the stack. In practice, these boundaries are not rigid, because the layers share and integrate hardware and software resources. This means many of the benchmarks cover many layers of the stack rather than just one. The middleware contains many fascinating challenges and poses an incredibly fertile area of innovation for engineers, scientists and mathematicians to solve the complex bottleneck problems needed to make FTQC a reality. Suitable benchmarking is critical here, highlighting important metrics of the early FTQC-era.
First, consider the compilation layer that turns the algorithms into hardware-executable circuits. Suitable benchmarks here are about the success rate and optimising the implementation of the algorithm. Some factors that affect the efficiency of using the compiled code on the quantum computer are the required circuit depth, number of logical qubits, types of gates needed, and the mapping and routing overhead needed for the type of computer the algorithm is run on. The compilation time is important when compiling is needed in real-time but here it is less of a concern. Useful benchmarks can be made by comparing the different requirements of a compiled code in different quantum architectures.
Second, the requirements needed to make the computation fault-tolerant are highlighted in the fault-tolerant orchestration layer. The number of quantum operations (QuOps) that can be reliably performed, including gate operations and measurements, is a key benchmark for FTQC. This also provides a useful indication of the robustness of the system and its capability for performing advanced computations.
Third, benchmarking the logical qubits and operations layer involves the number of logical qubits, logical clock rate, logical gate error rate and the Clifford volume12. The number of logical qubits the computer can maintain is currently one of the most used benchmarks by qubit hardware providers. The choices of qubit modality and error code directly affect this benchmark. However, this figure should be compared with the logical clock rate to understand the overall system’s usefulness. The logical clock rate is the rate at which fault-tolerant logical operations can be completed. The rate governs how fast the computer takes to solve a problem, affecting the resources needed to run the machine and the value of the solution. For instance, solving a real-time logistical problem such as directing tankers to different piers in minutes is more valuable than providing a solution in a day. This benchmark represents many different performance metrics about the middleware and the low-level qubit technology of a fully functioning FTQC. Consequently, the benchmark is not particularly useful today as there is no quantum computer with all the integrated parts to make the clock rate a meaningful figure of merit. The logical gate error rate is a useful metric as it shows whether the gates can be reliably operated. The Clifford volume is an early attempt to provide the quality of gate operations for an arbitrary system size and depth. This benchmark is restricted to only considering Clifford gates. These gates are readily available today, needing fewer qubits, and can be compared with other quantum computing architectures. Under the Gottesman–Knill theorem, quantum circuits composed only of Clifford operations can be efficiently simulated using classical computers. Consequently, these circuits do not provide universal quantum computational advantage. It is important to understand that commercially useful quantum computing will emerge with the implementation of the more complicated non-Clifford gate operations. This means this benchmark is important today in the early FTQC-era but will become less valuable as quantum computers evolve. Benchmarks involving the use of non-Clifford gates such as T-gate count, T-gate depth and the resources required for making non-Clifford gates will emerge as being more important.
Fourth, the quantum error correction (QEC) layer is responsible for detecting and interpreting physical-qubit errors, then updating the logical error frame to preserve reliable quantum computation. QEC is currently a focus for many organisations in resolving an important bottleneck. This layer has created five important benchmarks: QEC overhead, throughput, latency, logical error rate per QEC cycle and lambda, Λ. The QEC overhead represents how many physical qubits are needed to represent one logical qubit. This overhead is affected by the code being used (e.g. surface, qLDPC, color, etc), the fidelity of the physical qubits, the circuit depth and many other factors. Throughput determines whether the decoder can keep up with the syndrome stream of data coming from the qubit measurements. As the number of qubits increases the amount of data dramatically increases. If throughput is too low, syndrome data accumulates faster than it can be interpreted. This prevents the system from maintaining an up-to-date view of the logical qubits and may force the logical clock to slow down. If the decoder falls too far behind, errors accumulate faster than they can be tracked, corrected or incorporated into the logical frame, increasing the risk of logical failure. This is known as the backlog problem. Latency determines whether the QEC layer can return an updated logical frame or feed-forward decision quickly enough to support the running computation. The logical error rate per QEC cycle measures the absolute probability of logical failure during each correction cycle. One of the most important benchmarks emerging in the early FTQC era is Λ. This benchmark refers to the logical error suppression factor as code distance increases. A larger Λ is better. If Λ > 1, increasing code distance suppresses logical errors under the measured conditions. If Λ is one or below, scaling the code does not improve logical reliability. Importantly, this benchmark is measurable and conveys information about the key metrics of not just the QEC layer, the decoder’s algorithm and approximations, but also many layers of the middleware such as the code choice, syndrome extraction circuit and measurement quality. Also, Λ gives a sense of the quality of the physical qubits, physical qubit error rates, leakage and dephasing. Λ is a systems-level benchmark showing whether the hardware and middleware are working together to produce meaningful logical error suppression.
Two important benchmarks for fault-tolerant quantum computing are the logical error suppression factor, Λ, and the logical error rate per QEC cycle. Together, they show both whether quantum error correction improves as code distance increases and how reliable the logical qubit is under the reported operating conditions.
At this stage, the principal signal is that the reported values exceed one, although the precise values and experimental conditions remain important when comparing systems. The values obtained are from narrow operating conditions that do not reflect the values expected in a fully operational quantum computer. The fact that these values exceed one provides evidence of below-threshold logical-error suppression under the reported experimental conditions. The absence of Λ values from other leading organisations is not necessarily a sign of being behind technologically. Some companies will want to publish this benchmark when they develop a more mature system, perhaps involving Clifford and non-Clifford gates or operating an algorithm. However, in an industry that contains a lot of hype, this benchmark is useful for measuring technological progress. In addition to the reported Λ values, the associated logical error rates per QEC cycle provide interesting comparisons between modalities and architectures. Based on the experimental conditions for each reported value, further comparisons are left for the reader to interpret.
Finally, the control, readout and calibration layer requires its own benchmarks because it needs to operate the qubits in real-time. Relevant metrics include pulse fidelity, timing jitter, readout discrimination latency, feed-forward latency, measurement-induced crosstalk, calibration stability and parallel control scalability. In the FTQC era, the most important control benchmark may be whether the controller can integrate measurement, decoding and feed-forward quickly enough to sustain the logical clock rate.
The relative importance of benchmarks used to describe the quantum hardware components such as number of physical qubits and gate fidelities has changed in response to the different requirements of FTQC over NISQ. But the quality of the physical qubits is still paramount such as their energy-relaxation and dephasing coherence times (T1/T2). Improvements in the quality of physical qubits mean that the overheads required for FTQC are dramatically reduced. For instance, if topological or Majorana-based approaches deliver intrinsically protected qubits as intended, they could reduce the computational data demands on the middleware layer. Other relevant benchmarks are the single- and two-qubit gate fidelities and leakage rate.
Benchmarks create an evidence base from which standards can be written by establishing universal definitions and performance metrics for quantum computing. Recently, the UK government launched the National Quantum Standards Network (QSN), led by the National Physical Laboratory13. Work from this initiative should help identify which benchmarks are most useful, improve how results are reported and make comparisons between systems more reliable.
Today, there are several benchmarks that provide valuable insight about the usefulness and capabilities of early fault-tolerant quantum computers. These are time to solution, qubits to solution, the logical error suppression factor, Λ, and the logical error rate per QEC cycle.
In conclusion, benchmarking, when applied correctly, helps keep the industry rooted in evidence. It allows genuine progress to be recognised, weak claims to be challenged and hype to be separated from useful technological development. As quantum computing matures, benchmarking will increasingly influence where capital is invested, which technologies are adopted and ultimately which architectures survive.
Avoid the hype. Use meaningful, transparent benchmarks to demonstrate progress, identify bottlenecks and build confidence in the technology.
Use benchmarking to distinguish credible technical progress from hype and identify the companies most likely to deliver useful systems.
Use benchmarking to monitor industry progress, direct public investment towards removing critical bottlenecks, and support the development of trusted standards and measurement infrastructure.
Next move
Whether you need to understand, invest, partner, build, integrate, communicate or wait, Orbsight will help turn uncertainty into a practical route forward.
Images were generated by OpenAI ChatGPT using the author’s photographs and prompts.