Luminous fiber-optic network cables routing high-speed data into enterprise AI server racks, representing low-latency scale-up interconnect fabric.
Article Icon
Amy Reifenrath
@
TechArena
Aug 18, 2026

Astera Labs on Why the Fabric Decides AI Inference Costs

The cost of serving a model increasingly comes down to how fast accelerators can move data between one another as AI inference workloads shift toward multistep reasoning and mixture-of-experts models.  

Faster chips are not the answer, said Aanchal Sharma, senior director of product management at Astera Labs, because they buy more compute, not less waiting. Put the fastest chip in the world behind a slow fabric, and you have built what she calls “an expensive space heater.”

We sat down with Aanchal as part of our series about how companies that build AI infrastructure are seeing requirements change as deployments scale. Astera Labs builds fabric switches and connectivity products that let AI accelerators from different vendors work together inside a rack.  

She talked about what it takes to prove multi-vendor interoperability before deployment, why Astera Labs builds Scorpio and Taurus as open rather than proprietary interconnects, and where the fabric has to scale next as clusters grow toward hundreds of thousands of accelerators. Here’s what we learned.

Q1: Rack-scale AI infrastructure now depends on many vendors' silicon working together inside one system: accelerators, memory and fabric all from different suppliers. What does that interoperability require in practice, and where does it tend to break down?

A: At rack scale, physics doesn't care whose logo is on the silicon. A GPU from one vendor, a CPU from another, memory from a third: They all have to hold a stable low-latency link under real production traffic every hour of every day for years. That's the bar.  

Most failures show up exactly where you'd expect: signal integrity breaking down across longer or noisier channels, firmware that was never tested together choking on link training, and faults that only appear once hundreds of accelerators fire the same collective operation at the same instant.  

We stress-test these exact combinations in our Cloud-Scale Interop Lab before a customer ever racks a single unit so customers can deploy with confidence and focus their engineering resources on AI innovation rather than infrastructure integration challenges.

Q2: Fabric switches and retimers now have to work across GPUs, CPUs and memory sourced from different vendors within the same rack. What does proving that kind of multi-vendor interoperability involve before a customer can put it into production?

A: It involves validating the rack as a complete system, not just each component in isolation.  

We re-create the customer's topology with the specific CPUs, GPUs, memory devices, fabric switches and retimers, then test across operating systems, software stacks, workloads and protocols. Working with ecosystem partners, we use co-emulation, test automation and continuous regression testing to identify compatibility issues and confirm seamless integration before deployment.

The result is a validated configuration that reduces integration risk and accelerates time to deployment. That's the entire point of an interop lab: catch every integration risk on our bench months before a customer ever has to.

Q3: Astera Labs builds Scorpio and Taurus as open multi-vendor scale-up fabric rather than a single vendor's proprietary interconnect. What does an open fabric let operators do that a closed single-vendor interconnect can't?

A: An open fabric means an operator never has to bet the entire rack on one company's roadmap. Mix the best GPU with the best CPU with the best memory; swap a supplier when lead times blow out; carry hardware forward across generations instead of tearing it out.  

Taurus supports this model across Ethernet, UALink and ESUN, while Scorpio provides an open software-defined fabric architecture that supports diverse accelerators, system topologies, and both open and platform-specific protocols.  

Compared with a closed interconnect, this keeps architecture and supplier choices open as performance, availability and platform requirements change.

Q4: As inference moves into production, a growing share of what it costs to serve a model comes from moving data between accelerators, not from the accelerators themselves. Why can't faster chips alone fix that?

A: Faster chips buy more compute. They don't buy less waiting.  

As inference workloads move toward multistep reasoning and mixture-of-experts architectures, more of the total job becomes accelerators talking to each other rather than accelerators doing math. Put the fastest chip in the world behind a slow, high-latency fabric and you've built an expensive space heater, with GPUs sitting idle waiting on data instead of generating tokens.  

That's exactly why Scorpio builds acceleration for collective operations, like Hypercast, directly into the fabric; why Taurus keeps those links clean and low power across Ethernet, UALink and ESUN as clusters scale; and why COSMOS gives operators the visibility to find and kill communication bottlenecks in real time. Tokens per watt and tokens per dollar get decided in the fabric long before anyone reads a chip's spec sheet.

Q5: As clusters grow toward hundreds of thousands of accelerators spread across racks and rows, where does the connecting fabric have to go next? How do open standards and products built on them fit that path?

A: The fabric has to scale in three directions at once: scale up inside the rack, scale out across racks and rows, and increasingly scale across between clusters and sites.  

Scale-up keeps pushing toward higher radix and lower latency inside the rack. Scale-out has to connect racks and rows without the distance tax of added latency and cost eating the gains. Scale-across is the newest of the three, holding performance together across distances that scale-up and scale-out were never built to cover.  

That means more optical connectivity, higher-density switching, and fabric that carries intelligence, not just bits, accelerating collective operations instead of passively relaying them. Open standards like CXL, Ethernet, PCIe, UALink and ESUN give customers one way to build that path without betting the company, but that's only part of the story. Some of the largest deployments we support run on NVLink Fusion, and others need a fully custom link built around one customer's architecture. Products like Taurus and Scorpio are built to make all three paths work.

Subscribe to Our Newsletter

Read the latest in the world of AI, data center, and edge innovation.