A glowing chip rising above stacks of chips.
Article Icon
Amy Reifenrath
@
TechArena
Aug 14, 2026

Cirrascale on Matching AI Workloads to the Right Chip

We sat down with David Driggers, CEO and founder of Cirrascale Cloud Services, a neocloud that runs dedicated hardware for AI workloads across NVIDIA, AMD, Qualcomm and Tenstorrent accelerators. Cirrascale’s approach is relevant now as enterprises move from proof-of-concept AI deployments toward production and need infrastructure that matches each workload to the chip built for it.

We talked about how Cirrascale matches workloads to the right accelerator as models grow, how its billing model creates cost predictability, and what separates neoclouds built for low-latency inference from those built for training alone. Here's what we learned.

Q1: Why run every accelerator on dedicated hardware?

A: Cirrascale gets classified as a neocloud, but we were building this hardware before the category had a name. We designed the industry’s first 8-GPU server back in 2012 working directly with NVIDIA, and we were infrastructure partners to OpenAI when they were still an 8-10 person team. That hardware DNA is why we can run NVIDIA, AMD, Qualcomm and Tenstorrent side by side, alongside our work with Google on Google Distributed Cloud, and actually know what each one is good for. No single chip serves a 1B-parameter model and a 400B-parameter model economically, so our customers get the accelerator that fits their workload instead of bending their workload to fit whatever one vendor is selling.

Q2: How does billing create predictability?

A: We don’t charge for data ingress or egress, which removes one of the biggest hidden cost swings customers deal with elsewhere. We also give customers the choice between token-based pricing for intermittent workloads and GPU-hour billing for anything running 24/7, so the bill follows how they use the platform. And because we track token consumption at a granular level, customers get real insight into where they’re burning input tokens inefficiently. That data helps them optimize the model, not just the invoice. Inference is forever, as we like to put it, and those costs compound. That’s why we built a cost calculator: Customers can model their own workload before they commit to anything.

Q3: What’s standing between customers and more capacity, and what’s changed?

A: Most enterprises are still in the POC phase with models like Gemini, GPT and Claude, trying to prove utilization before they can get more approved budget. Nobody’s handing out incremental AI spend without proof it moves the needle. That’s part of why we run Private Gemini on Google Distributed Cloud. It gets a frontier model into production inside the customer’s own security boundary instead of leaving it stalled in evaluation. On the hardware side, we’re pushing into multimodal, larger-context inference. What’s changed in the last year is that hybrid deployment, on-prem for steady workloads plus neocloud for real-time workloads across regions, is becoming the norm. I expect that to be standard by the end of this year and accelerate hard through 2027.

Q4: What have you learned matching workload to chip?

A: We built the load balancing in the Cirrascale Inference Platform ourselves, specifically because a model’s ideal hardware changes as it grows. An 8B model that becomes a 70B model isn’t cost-optimal on its original chip anymore, and that kind of migration is really hard to do on-prem. The single biggest cost lever is capacity utilization. The platform puts each workload on the most efficient capacity that still clears its latency target, so customers aren’t paying premium rates for headroom they don’t need.  

Q5: What separates the neoclouds that thrive from the ones that fade?

A: Training and inference are completely different animals. In training, you can run a thousand GPUs and barely notice a network hiccup. In inference, that same hiccup is a failed request in front of a customer. The neoclouds that survive are the ones built for real-time workloads and low latency from the ground up, not the ones bolting inference onto training infrastructure. Carrier-hotel-grade connectivity is a big part of that, which is why we partnered with Telehouse to deploy the Cirrascale Inference Platform inside their data centers. It also means being able to serve regulated industries, where FedRAMP High and CMMC 2.0 are the price of entry rather than a roadmap item. We also give customers real ownership flexibility. Some customers own their GPUs outright while we run the networking and storage around them. And we’re expanding our data center footprint this year to put that capacity and failover closer to where the inference actually happens.

Subscribe to Our Newsletter

Read the latest in the world of AI, data center, and edge innovation.