Extreme closeup of fibers weaving together into a fabric.
Article Icon
Amy Reifenrath
@
TechArena
Aug 11, 2026

Hedgehog on Turning AI Network Design Into a Product

As the AI Infra Summit approaches, TechArena has been chatting with the builders of AI infrastructure about how their requirements are changing as deployments scale.

We sat down with Marc Austin, CEO and co-founder of Hedgehog, a company that builds open software-defined networking for AI clusters. Hedgehog’s approach is mattering more now, as AI infrastructure spreads from single racks to multiple sites and a single misconfigured switch can leave GPUs across a cluster waiting on the network.

We talked about how Hedgehog’s software replaces switch-by-switch network engineering with a declarative operating model, why the company holds reference architecture validation from both NVIDIA and the Open Compute Project, and what it takes to keep job completion times down as clusters scale across sites. Here’s what we learned.

Q1: Hedgehog builds an open software-defined fabric that operators manage as one system rather than as a network stitched together switch by switch. What does that let an operator do at scale that the box-by-box model can’t?

A: It turns the network from a monthslong engineering project into a product you deploy.  

Designing an AI network from scratch means making thousands of interdependent decisions (topology, QoS, congestion control, failure domains) and then keeping them consistent across hundreds of switches by hand.  

With Hedgehog, an operator declares the cluster they want, and the fabric handles Day-0 through Day-N automatically: provisioning, lossless Ethernet tuning, congestion-aware routing and failure recovery, all through Kubernetes-native operations.  

That’s what lets a lean cloud team network like a hyperscaler.  

Q2: Hedgehog supports the NVIDIA Cloud Partner reference architecture, validated for Spectrum-X, and the Open Compute Project (OCP) reference architectures for AI training and inference. Why does that dual validation matter to an operator deciding how to build an AI cluster today?

A: At NVIDIA GTC 2026, we announced support for NVIDIA Spectrum-X Ethernet and the NVIDIA Cloud Partner reference architecture. At the 2026 OCP EMEA Summit, our AI training and inference fabric designs became OCP Accepted: 100% open source, hyperscaler reviewed, with complete BOMs available today on the OCP Marketplace.  

An operator can standardize on Spectrum-X, deploy a fully open SONiC-based fabric on OCP hardware, or run both, and keep a single declarative operating model across all of it. No other network vendor gives operators that freedom at the architecture layer.

Q3: In an AI cluster, every idle GPU is usually waiting on the network. What does the fabric have to get right to keep job completion times down, and why is that harder to do with closed, manually operated networks?

A: The metric that matters is job completion time, and the network defends it or destroys it.  

The fabric has to get five things right at once: lossless Ethernet with properly configured RoCEv2 QoS; congestion-aware routing; rail-optimized topology; dual-plane resilience, so a single failure never stalls a training run; and telemetry with automated recovery, so problems are fixed faster than a GPU can starve.  

In a closed, manually operated network, each of those is a specialist tuning exercise repeated per deployment. Any one misconfigured switch quietly taxes every job on the cluster. Hedgehog ships them as validated defaults in the reference architecture and enforces them continuously in software so that the fabric stays out of the GPUs’ way.

Q4: As infrastructure becomes composable, the network becomes the reassembling component. What does that demand from network software?

A: Composability is fundamentally a network abstraction problem, and hyperscalers already showed the answer: open networking plus VPC abstractions. Hedgehog brings that same model to AI clusters.  

Tenants get VPCs that can span multiple accelerator types and pods, with policy-based multitenant isolation, composed and recomposed in software rather than recabled. We’ve proven this live with disaggregated inference: NVIDIA GPUs handling prefill, SambaNova RDUs handling decode and Intel Xeon orchestration, all running across one Hedgehog fabric with no network penalty.  

Because the whole thing is operated through a declarative, Kubernetes-native model, AI infrastructure can be run with a small team.

Q5: As clusters grow, multiply, and spread across sites, where does network software have to go next? What role do open reference architectures play in letting the industry scale without every operator redesigning the network from scratch?

A: Network software has to scale in two directions at once: up and out.  

“Up” means prescriptive scale-out designs, with our OCP reference architectures defining Open Pod Group scalable units from 64 to 1,024 xPUs and a roadmap to Ultra Ethernet as the interconnect evolves.  

“Out” means treating many pods, clusters, and sites as one operable system, with hybrid multicloud routing and multitenant security built in rather than bolted on.  

Open reference architectures are what make that repeatable: Instead of every operator rederiving the design math, validated architectures with complete BOMs are published on the OCP Marketplace for anyone to build from. Everyone gets to network like a hyperscaler, and the industry scales without redesigning the network from scratch each time.

Subscribe to Our Newsletter

Read the latest in the world of AI, data center, and edge innovation.