Supercomputing News logoSupercomputing News logoBeta
AIHPCQuantumEmerging
Subscribe
Supercomputing News logoSupercomputing News logo
Pillars
AI—HPC—Quantum—Emerging—
Theme
Subscribe
Supercomputing News logoSupercomputing News logo

Trusted reporting on AI, HPC, Quantum, and the technologies shaping the future of computing. Cryptographically signed. Agent-accessible.

Pillars

  • Artificial Intelligence
  • High-Performance Computing
  • Quantum Computing
  • Emerging Technology

Entities

  • Organizations
  • Products
  • People
  • Places

Publication

  • About
  • Contributors
  • Topics
  • Contact
  • For Agents

Weekly Update

Keep track of the biggest stories in supercomputing, every Thursday.

Subscribe for free today
© 2026 Supercomputing News
Privacy PolicyTerms of Use
Quantum ComputingQuantumNews

Quandela and NVIDIA Draw the Control Boundary for a Photonic QPU on NVQLink

The joint white paper says the GPU never controls the quantum process and quotes an error-correction budget Quandela's photonic roadmap does not yet reach.

Top-down illustration: a cyan line divides a grid of GPU compute blocks from a photonic chip, with one thin link ending at a controller.
One link crosses the boundary and stops at the controller. In the Quandela-NVIDIA NVQLink architecture, the GPU joins the loop while Quandela's FPGA-based Quantum System Controller keeps control of the photonic QPU.AI-generated / SCN
SCN Staff
The Squad
Published
Sep 17, 2026
Add Supercomputing News as a preferred source on Google
Reading0%
Listen to this article14 min
Loading audio…
0:00/ 13:38PausedMuted
Played in full
Audio unavailable
Speed
1×
Download audioMP3 · 12.5 MB
0:00

On September 14, 2026, Quandela and NVIDIA published a joint technical white paper, AI-Native Hybrid Computing Architectures: From Algorithm Discovery to NVIDIA NVQLink Integration and Scale, describing a CPU-GPU-QPU architecture in which a Quandela photonic quantum processor is attached to NVIDIA GPU infrastructure over NVQLink. The same day, NVIDIA's newsroom announced CUDA-Q Logical, which NVIDIA describes as an orchestration layer meant to give developers a programmable, verifiable way to build applications for fault-tolerant quantum computers, and named three companies that had integrated with NVQLink: Anyon Computing, Quandela and Quantum Machines. The announcements coincide with IEEE Quantum Week in Toronto, which runs September 13 to 18. Quandela says it will present the architecture at NVIDIA's booth, number 600, on Thursday morning, September 17.

In April, Supercomputing News argued that NVIDIA's Ising model launch was really about quantum computing's classical control plane, and said the next proof point to watch would be evidence that vendors could integrate those models into live control workflows without breaking latency, cost, or operational complexity. The Quandela paper is not that proof point for the Ising models themselves; it mentions them once, as a model family that complements the stack for calibration and decoding. What it does supply is an unusually detailed account of where the GPU's job ends and the quantum controller's begins. Read alongside its own references, it also shows how far the current demonstrator sits from the error-correction workload the link was designed for.

Three processors, one boundary

The paper divides labor in plain terms. The CPU is the host and orchestrator; the GPU does the dense parallel computation and handles most of the computation around quantum execution; the QPU contributes the specific kernels where a quantum method is expected to pay off. In this model, simulation, digital twins of the device, variational parameter updates, and error-correction decoding are all GPU work, expressed in CUDA-Q as a single program whose quantum kernels can be retargeted from a simulator to a physical processor.

Weekly Update

The biggest stories in supercomputing, once a week.

AI, HPC, quantum, and emerging tech. Reported, not aggregated.

Free · no account · unsubscribe anytime

The sentence that matters for any QPU vendor weighing this stack is on page 17: "GPU–QPU integration does not imply that the GPU directly controls the physical quantum process. The QPU retains its dedicated control stack, including the electronics and real-time functions required to operate the photonic processor." On Quandela's side, that control stack is the Quantum System Controller, or QSC, where FPGA-based control electronics handle the deterministic, real-time functions the photonic QPU needs. The GPU side runs CUDA-Q Realtime, NVIDIA's runtime for exposing GPU functions to a quantum controller as callbacks with microsecond latency rather than as jobs submitted over an HTTP interface. The paper states the intent in one line: cut the communication and synchronization overhead between the two computational domains, and leave quantum control where it is.

The transport is RoCE, RDMA over converged Ethernet, which the paper calls the basis of NVQLink's open reference architecture. NVIDIA uses "open system architecture" in the launch release and "platform architecture" on the product page. In the NVIDIA product and release pages cited here, NVQLink is described as an architecture or a platform; neither page uses the word standard.

The session model comes from Quandela's June engineering note, reference [5] in the paper. The CPU reserves the QPU, preconfigures the photonic chip, and hands control to the QSC, which the note says becomes the sole owner of the quantum resource for the duration of execution. Existing HPC schedulers, according to Quandela's June release, keep responsibility for reservation, allocation and accounting. That is the pattern Supercomputing News described last week at Jülich, RIKEN and Hartree: the QPU as an accelerator the workload manager allocates, with a private low-latency path once the session is open. One footnote to the paper's reference, on page 14, on scheduling by the same workload manager: NVIDIA has owned Slurm's developer, SchedMD, since December 2025.

Quandela's June note adds something the white paper does not spell out. The note describes Quandela's use of the link as the opposite direction of request from NVIDIA's usual framing. In NVIDIA's reference design, the GPU is a compute resource the controller calls on, typically to decode error syndromes; in Quandela's demonstrator, the GPU calls on the QPU, consuming it as an accelerator inside an inference pass for quantum machine learning.

Whose numbers

Every latency figure in this story has an owner, and the paper mixes them.

A round trip of under four microseconds is NVIDIA's NVQLink specification. The product page lists less than 4.0 microseconds for the FPGA-to-GPU-to-FPGA round trip, 400 Gb/s of bandwidth, and 40 PFLOPS of FP4 compute. NVIDIA's reference measurement, published in a November 2025 technical post and in the arXiv architecture paper the white paper cites as reference [17], is a mean of 3.84 microseconds and a maximum of 3.96 microseconds on an Arm host with an RTX PRO 6000 Blackwell GPU, a ConnectX-7 network card, and an AMD RFSoC FPGA. None of that is Quandela hardware.

The paper is inconsistent with itself in this figure. Page 15 gives "round-trip latency below four microseconds."

Page 16, citing the same NVIDIA source, describes "an RDMA-on-Ethernet link with sub-microsecond round trips." NVIDIA's published specification and measurements support the first phrasing; nothing in NVIDIA's public record supports the second. Supercomputing News has not confirmed with Quandela which figure was intended.

The paper publishes no measured latency for the Quandela link. Quandela's June release says the company validated the path by measuring low-latency communication between the GPU infrastructure and the FPGA-based Quantum System Controller, and gives no figure. The June engineering note describes the NVQLink networking as sub-millisecond. The only Quandela-specific number in circulation is spoken: Sam Stanwyck, NVIDIA's director of quantum product, told Laser Focus World in August that Quandela "recently demonstrated an NVQLink that reduced their latency by an impressive 200x," with no baseline or method stated. For comparison, Xanadu, another photonic vendor, published end-to-end classical-quantum loop times of under 3 microseconds for its Backline platform with AMD on September 10, built on PennyLane rather than CUDA-Q.

The numbers Quandela does publish are on the QPU side. For current-generation Quandela hardware, its June note gives about 20 milliseconds for an incremental phase update, about 10 milliseconds to acquire roughly 1,000 samples, and about 30 milliseconds of total quantum processing per datapoint, against at least 5 seconds per datapoint through what the note calls a conventional cloud quantum stack. That demonstrator ran on a DGX Spark, NVIDIA's GB10-based desktop system, connected through a ConnectX-7 card to the FPGA-based QSC. The white paper describes the hardware only as NVIDIA GPU infrastructure, and does not say whether the September configuration differs.

Analysis: for the workload Quandela documents, a microsecond link is not where the time goes. The QPU-side step is tens of milliseconds per datapoint. NVIDIA's specification puts the link's round trip under four microseconds, roughly four orders of magnitude smaller, and that figure is the specification, not a Quandela measurement. Quandela sets at least five seconds of conventional cloud-path latency beside about 30 milliseconds of QPU processing, but it publishes no end-to-end latency for the integrated path. The June note says that path bypasses the host CPU, the Python runtime, REST APIs, remote schedulers, and serialization layers, and the paper frames the objective as lower communication and synchronization overhead, so removing those layers is the design goal. The paper has not published how much of the five seconds the integrated path removes in practice. The microsecond budget exists for a different class of workload, which is where the paper's error-correction numbers come in.

The budget the paper quotes, and the roadmap it points to

The paper is precise about why tight coupling matters, and it uses NVIDIA's numbers to say so. In a QPU running an error-corrected program, it says, stabilizer measurements emit a continuous syndrome stream at cycle rates approaching 1 MHz, and decoding that stream while the encoded state remains valid is estimated to require network throughput of order 1 Tb/s and compute of order 1 PFLOP/s, within latency tolerances of tens of microseconds. The citation is reference [17], Caldwell et al., NVIDIA's own architecture paper, which gives a syndrome rate of up to about 1 MHz, throughput of about 1 Tb/s, and compute of up to about 1 PFLOP/s. Those are estimates for a fault-tolerant machine in general. They are not measurements from any Quandela system, and the paper does not present them as such.

The paper then places fault tolerance in the future. Its three-phase table runs Access (Perceval, MosaiQ, Quandela Cloud), Integration & Discovery (MerLin, GPU simulation, CUDA-Q simulation), and Scale, where the listed enablers are NVQLink for architecture validation, leading on to SPOQC with CUDA-Q support for long-term scaling. SPOQC is Quandela's spin-optical quantum computing architecture, published in Quantum in July 2024 and described as an adaptable, modular hybrid architecture for fault-tolerant quantum computing that combines quantum emitters with linear-optical entangling gates. The white paper says future SPOQC architectures build on the same software and hybrid execution principles while progressively introducing logical qubits, deterministic entanglement, and fault-tolerant operation. No date appears in the paper. Quandela's roadmap page says "large-scale fault-tolerant quantum computers by 2028 and beyond," the same year the US Department of Energy has set as a target that Supercomputing News has reported is hard to define across vendor roadmaps.

So the demonstrator is an integration architecture aimed at variational and quantum-machine-learning workflows, and the error-correction case that would consume a microsecond loop is not demonstrated on Quandela hardware in the paper. The paper does not claim otherwise; it frames this as forward compatibility: build the loop now so SPOQC can use it later. NVIDIA's real-time decoding demonstrations to date have run on Quantinuum's trapped-ion systems, including the GH200 qLDPC decode SCN noted in August; the wider decoder landscape, where syndromes arrive at microsecond intervals and the decoder must keep up, is the workload that sets NVQLink's budget.

Nothing in the paper is a deployment. No center, customer, or system is named; Quandela's June release, for its part, refers to future NVQLink-enabled MosaiQ deployments. Quandela was one of 17 QPU builders NVIDIA named when it launched NVQLink in October 2025; it does not appear on NVIDIA's November 2025 list of supercomputing centers adopting NVQLink because that list includes centers rather than vendors, nor among the partners named at the CUDA-Q Realtime launch in March 2026.

The other two integrators

Quantum Machines, a control-system vendor, said on September 13 that it had become "the first control company to run an end-to-end NVIDIA CUDA-Q program across live qubits and a PPU classical processor with NVIDIA NVQLink," with the full exchange completing in about a millionth of a second, and that it is demonstrating the run live in Toronto. The release, carried by The Quantum Insider, does not name the qubit hardware; Quantum Machines' own site did not serve the release when checked. Anyon Computing, a superconducting QPU builder, announced an open-source real-time control system integrated with NVQLink over an RDMA-over-Ethernet fabric, with its QPU installed inside a commercial AI data center paired with NVIDIA GPUs, according to Quantum Computing Report; Supercomputing News could not locate Anyon's own September announcement. The deployment itself is on Anyon's record. In a March release, the company said SDT Inc. had launched what Anyon called the first commercial tightly coupled hybrid quantum-classical data center, in Korea, with Anyon's QPU connected to NVIDIA accelerated computing over NVQLink.

Analysis: three integrators across three roles: a controller vendor, a superconducting QPU builder, and a photonic QPU builder. What the April piece called the control-plane strategy now has partner documents behind it. Anyon's has a commercial site on record. Quandela's is the one that writes down the boundary condition, with the GPU as a participant in the loop and never its owner, and that sentence is what lets a QPU vendor adopt the stack without ceding its control electronics.

Sovereignty, stated as fact

Quandela is French. Its June release named HPC centers, sovereign AI and quantum programs, advanced research organizations, and industrial users as the audience for a deployment model in which, the release says, a customer-owned photonic QPU could be installed on premises. GENCI, France's national supercomputing agency, was on NVIDIA's November 2025 list of centers adopting NVQLink. No published document connects the two, and SCN has reported that on-premises quantum procurements vary widely in what the buyer owns.

What would move this

Four things would change the record: a measured round-trip figure for the QSC-to-GPU path published by Quandela or NVIDIA; a correction of the latency the paper cites on page 16; a named center with a date; or a logical-qubit demonstration on Quandela hardware with a SPOQC timeline attached. Until one of those appears, the paper is an architecture statement with an explicit division of control, an integration demonstrator on DGX Spark aimed at quantum-machine-learning workflows, and an error-correction budget borrowed from the vendor that wrote the specification. Quandela's booth session in Toronto is Thursday morning.

Quantum Classical Control PlanePhotonic QuantumNVIDIAQuantum Error Correction
AI disclosure
This article was prepared with AI assistance for research and drafting under human direction and editorial control, per SCN house style.
About the contributor
SCN Staff
The Squad

The SCN Staff is a small AI editorial squad working under human direction. Each agent owns one job.

Scout does the research. It runs down primary sources and checks what's already been published, on SCN and everywhere else, before a story gets written. If a claim can't be traced back to a real document, Scout flags it.

Forge writes. It takes what Scout found and turns it into a draft, argument and sentences and all. Every SCN piece starts here, then gets sharpened.

Cipher handles search: the titles, descriptions, and keyphrase work that decides whether a good article ever gets found. Least glamorous job on the squad. Also one that matters more than it looks.

Pixel makes the visuals. Images, charts, the occasional diagram, all built to SCN's brand instead of pulled from a stock library. When something's easier to see than to read, it goes to Pixel.

Editorial judgment and the final call stay with the humans. So does the fact-checking.

Related reading
Quantum · NewsNVIDIA’s Ising Pitch Is Really About Quantum’s Classical Control PlaneQuantum · AnalysisJülich, RIKEN and Hartree Show Different Stages of Quantum-Supercomputer IntegrationAI · AnalysisThe CUDA Moat Is Becoming a Composability Problem