Cirrascale absorbs the porting for AMD, Qualcomm, and Tenstorrent silicon. Driggers's depreciation logic implies an older, written-down GPU is hard to beat.

Cirrascale Cloud Services moved its Inference Platform from preview to production on September 15, 2026, 18 months after its early preview at Nvidia GTC on March 18, 2025. The pitch is unchanged: a serverless endpoint that chooses the accelerator for you. What changed is the accelerator list: the preview launched on Blackwell alone, and the production release routes across NVIDIA, AMD, Qualcomm, and Tenstorrent hardware "with no code changes required to switch."
In July, SCN argued that inference-chip challengers rarely die because of the chip; software maturity and volume economics do the killing. Cirrascale claims a neutral cloud can take the software problem off the buyer's desk. SCN sat down with CEO and co-founder Dave Driggers at the AI Infra Summit in Santa Clara to ask what the platform absorbs, what it does not, and which silicon gets the traffic.
The answers were more specific than the release, and in one respect they complicate it. The platform was built to make alternative silicon consumable, but its cost model, as Driggers described it, favors hardware further along its depreciation schedule. We used a second-life H100 as the example; on the arithmetic he laid out, an older Nvidia GPU is a plausible winner in many routing decisions. He gave no vendor split.
Cirrascale does the porting, but only for models it has already validated. Driggers said the company forks vLLM, the open-source serving engine, and tunes a version for each hardware type and usage pattern. The company tests a model against each accelerator before offering it there, measuring token rate, latency, and sustainable batch sizes. If a model hasn't been validated on a given vendor, the router doesn't fall back or improvise. "It won't even try to run the model where it can't run," he said.
The customer never sees that table. The buyer selects a service level, Driggers said, and the platform handles the rest: quantization has a standard setting the customer can raise or lower, batch sizes shrink for near-real-time work and grow for offline jobs, and placement is invisible. Fine-tuning typically happens on NVIDIA hardware because NVIDIA has the better software ecosystem for that job. Once a model is fine-tuned, it can run on whichever validated hardware Cirrascale chooses.
The abstraction layer is a curated catalog: Cirrascale removes the porting burden by doing the porting itself, per model, per vendor, and coverage of non-NVIDIA silicon is whatever that work has reached. The customer buys an SLA; whether it is met on a Qualcomm AI 100 or an H200 is Cirrascale's business, though the console shows management which models ran where, for how long, and at what utilization, with job codes for chargebacks.
That is a coherent answer to the software-ceiling question, and a narrower one than "no code changes" implies. SCN's August analysis of ROCm argued that closing the gap with CUDA needs test hardware more than it needs missing APIs; a cloud that validates every model on every vendor is running that test bench as a business. Driggers did not say how much of the open-model catalog that work has covered on each vendor.
Routing runs in a fixed order. Availability comes first: if the cheapest hardware that meets the SLA is busy, the job goes to something more expensive that is free. Cost comes second, and the customer's price does not move with the hardware. "We hold the price, and both eat and make up the variance," Driggers said. The customer agrees to a cost per token at an SLA; Cirrascale tries to hit it on the least expensive qualifying silicon, and pockets or absorbs the difference.
That cost calculation includes where a machine sits in its depreciation schedule.
"There's a real cost and a fake cost," Driggers said. On a five-year depreciation schedule, hardware is very cheap on the books after three years. A newer accelerator, even a faster one, starts its own schedule and carries a higher accounting cost, and the router accounts for that. He named Tenstorrent as the expensive example. It is a position he has held for a while: on theCUBE in October 2025, he said Cirrascale builds so that training hardware can be repurposed into inference, the equipment's second life and long tail.
The router and the porting are real, and the cost input Driggers chose to explain rewards whichever machine has been depreciating the longest. He did not share a traffic split in the interview. An H100 that entered service in 2023 is three years into a five-year schedule, the point at which, by his description, hardware is very cheap on the books. Cirrascale said Galaxy Blackhole entered broad commercial deployment on its cloud in May 2026, placing that hardware near the start of its own schedule. If the router weighs book cost as he described, the older NVIDIA hardware begins most comparisons with an advantage a newer chip has to overcome on real cost per token. For the challengers, that means validation gets a chip into the routing table; traffic follows as its accounting cost falls, or when its throughput beats a written-down GPU regardless.
Cirrascale says the production router supports four vendors. One announced system is not yet carrying traffic: AMD's Helios rack.
Public records show what Cirrascale sells as dedicated systems, which differs from what the router dispatches to. Qualcomm Cloud AI 100 systems have been on the company's price list for some time, alongside AMD MI300X and MI250 and NVIDIA hardware from the A100 to the B200. Tenstorrent Galaxy Blackhole went into what the company called broad commercial deployment on its cloud on May 4, 2026; the price list still says to ask.
Helios, with the MI455X, is not in production on the platform. Cirrascale announced support for it on July 14, 2026, with capacity to follow as the AMD platform reached general availability. Driggers put the delay on data center space rather than on AMD or Cirrascale. "It's 8,000 pounds. It's double wide. It's water-cooled," he said, and the large, modern facilities that can take it are booked until the first quarter of 2027, by his account. AMD's Helios page confirms the double-wide Open Rack Wide form factor and liquid cooling but publishes no rack weight; ServeTheHome put the CES 2026 unit at nearly 7,000 pounds. Either way, a rack holding 72 GPUs and 31 TB of HBM4 is a facilities project.
Cerebras, a Cirrascale partner since 2022, is no longer a standard offering. Driggers said Cerebras chose to sell its own inference as a service, leaving Cirrascale little need to resell it, though the two still cooperate when a customer wants a combination; a wafer-scale box has to be fed continuously to pay for itself, he said, and Cirrascale's service is built to scale up and down. He expects to announce two or three more accelerator vendors in the first quarter of 2027 and did not name them, saying he wants real benchmarks and true general availability first. The candidates are increasingly selling systems: d-Matrix's April acquisition of GigaIO's data center business was a bet that inference silicon sells as a rack, the form Cirrascale has to find floor space for.
Asked for production numbers from the preview period, Driggers gave one, about traffic rather than latency or cost. When the platform started, he said, requests ran at roughly two tokens in for every token out. Agentic workloads pushed that toward a hundred in for one out. He offered those ratios as characterizations, not audited figures; the shift, he said, forced Cirrascale to rework caching and hardware balancing and delayed the release.
He dated the turning point to about a month after OpenClaw, the open-source agent framework, took off. Cirrascale added what it calls agent guard shortly after, and Driggers's illustration came from inside the company: its president set an unguarded agent running and burned through about $200 in tokens overnight. The production release ships with spend controls across teams, tool-call allow-lists, per-agent caps and audit trails, plus what Driggers described as hardened agents that restart and finish when something fails mid-run. He tied that to economics as much as safety: the cheapest tokens are at night, when daytime real-time hardware is idle, and that is when he wants agents working.
CTO Alex Nataros, who took over the role from Driggers in February 2026, is quoted in the release saying enterprises no longer struggle to stand up a model endpoint; they struggle with governance, cost control and getting a secure application in front of employees. The one reference customer Driggers put on the record is Ai2, the Allen Institute for AI, whose OLMo, Molmo and Tülu models have been served on the platform since July 2025.
The sovereignty dimension of this release runs through one product. Cirrascale announced Gemini on Google Distributed Cloud on April 22, 2026, and began previewing it in connected and fully air-gapped configurations; the September release folds it into the platform's closed-model tier. The appliance is a Google-certified, Dell-built server with eight Nvidia GPUs, B200 today and B300 going forward. Driggers described it as a server Cirrascale owns and maintains, placed in its own data centers or at the customer's site; VentureBeat's launch report said either Cirrascale or the customer can own the box. Prompts stay in it. For patching, the appliance connects to Google, updates, and disconnects; nobody, including Cirrascale, gets inside the operating system.
That isolation created the software work. Google delivers the model inside the appliance and no tooling around it, Driggers said, because Google Cloud's management tools run on TPUs. Cirrascale built its own load balancer for multi-appliance deployments and its own user tiers, giving knowledge workers, advanced users and super users different token rates and priority. The model router sits on top: a request Gemma can handle does not go to Gemini, and a request that needs corporate data goes to the fine-tuned model.
Isolation is also how Cirrascale approaches compliance. Driggers said the inference service assigns a server to one tenant; before reuse, Cirrascale tears it down to bare metal and reprovisions it. He linked that isolation to the company's ability to handle HIPAA-regulated work. The release's "aligned with HIPAA, SOC 2, and FedRAMP requirements where required" is the accurate wording: Cirrascale is not FedRAMP authorized today.
On FedRAMP High, which he told TechArena in August was "the price of entry rather than a roadmap item," the status is in progress, with a target before the end of 2026. Cirrascale's Government Services division, launched March 10, 2026 with a Google Public Sector partnership, is the buyer this is aimed at, and he said all four vendors' silicon is available in air-gapped and on-premises deployments.
Driggers's case for why this matters is data-sovereignty regulation and the geography of inference: financial services firms, he said, keep the crown jewels on-premises and need a provider in the middle to run inference in every jurisdiction where the hyperscalers are not, and he expects sovereignty rules to multiply.
The release describes a platform where the buyer chooses an SLA and the silicon is somebody else's problem. The interview confirms that and adds the mechanics: validated models only, availability before cost, price held per SLA, and a cost model that counts depreciation. Cirrascale gives the alternative-silicon vendors a serving stack that has already been ported and tuned, something Qualcomm and Tenstorrent have struggled to get from enterprise buyers directly. Traffic is decided separately, by a router that, on Driggers's description, treats book cost as a real input. By SCN's reading, that gives the oldest NVIDIA hardware in the fleet a head start the interview neither confirmed nor quantified.
The next two quarters hold a few checkpoints. Driggers expects Helios production in the first quarter of 2027, when he says suitable data center space becomes available, and he expects to announce two or three more accelerators in the same quarter. FedRAMP High authorization is in progress, with a target before the end of 2026, according to him. The number that would settle the routing question- the share of requests landing on non-NVIDIA silicon- is one Driggers did not disclose in the interview.