img.jpg)
The world tends to measure artificial intelligence (AI) progress in two units: the number of graphics processing units (GPUs) deployed and the megawatts they consume. Yet the people who actually build these systems know that the GPU count is only one part of a much larger engineering effort. Power delivery, cooling, networking, and storage all have to come together before a single accelerator does useful work, and the discipline of fitting those pieces together is becoming one of the most consequential roles in the data center.
On a recent TechArena Data Insights episode, Solidigm’s Jeneice Wnorowski and I explored that hidden layer of the AI buildout with Hitesh Kumar, a GPU cluster architect at Nebius, an AI cloud company that has been drawing attention across the neocloud landscape. Our conversation moved from custom hardware and proprietary software to power heavy AI workloads.
Hitesh described his role as starting once a site has been selected and power secured, and ending when responsibility passes to the logistics and deployment teams. That gap, he explained, is where his team plans clusters, maps them onto floor plans, and decides what connects to what. The work is deliberately broad, and much of it has little to do with the accelerators themselves.
“My job as a GPU cluster architect focuses on a lot of things that aren’t GPUs,” he said. Power and cooling are a major part of the discussion, and so are networking, storage, and the management plane that orchestrates everything. In fact, it’s only after that supporting infrastructure is planned that the headline GPU count comes back into the picture.
Storage has always mattered for training AI models, as data must reach the GPUs fast enough to avoid stalls. Hitesh noted that inference, or running those models in production, is now creating fresh demand beyond that baseline. With a large language model (LLM) chat bot for example, the full state of a long conversation increasingly needs to move off the GPU, and sometimes off the server entirely, with storage acting as an intermediate tier between what a GPU holds close and what it will need soon.
Interconnects are evolving in a similar trajectory. Hitesh traced a path from today’s pluggable optics that enable a connection to one to two cables toward denser designs with tens of ports per unit, and eventually toward co-packaged optics that place optical engines directly in the server. Each step raises new challenges in cooling, cabling, and supply chain readiness that operators must manage alongside the mature technology they rely on today.
Adding GPUs sounds simple, but Hitesh pointed to two realities that catch teams off guard. The first is component failure. Citing Meta’s published study of a 16,000 GPU cluster, he noted that Meta’s team experienced a failure every three hours on average. At that rate, software must become genuinely fault tolerant, and operations, spares, and logistics all have to scale to match.
The second unanticipated challenge is power behavior. Hitesh highlighted the scale of the power that clusters now draw, and he described how a checkpoint pause can drop a rack tens of kilowatts in moments. Swings that large can affect the grid, pushing operators to consider capacitors, batteries, or software mitigations that rarely make the headlines.
For organizations early in their AI journey, Hitesh offered valuable guidance. Going from zero to one, he said, “you really don’t want to be thinking about buying your own infrastructure,” given the high startup costs and overhead. Cloud resources make sense first, followed by colocation, and eventually a dedicated site once demand justifies it.
Looking forward, he expects steady efficiency gains, with “every part of your stack” delivering more performance per dollar. The change he finds most interesting is in network topology. He anticipates clusters organized as “lots of small islands of very tightly connected GPUs,” linked by sparser scale-out and scale-across fabrics. He sees the same fractal pattern emerging across NVIDIA rack-scale designs, Google’s tensor processing units (TPUs), and Huawei’s accelerators alike.
Hitesh’s perspective provides a useful counterpoint to conversations that reduce AI infrastructure to a single metric. The teams that succeed will treat power, cooling, storage, and networking as first-class design decisions rather than afterthoughts, and will scale their operational maturity in step with their hardware. As models grow and inference proliferates, the competitive advantage will belong to operators who understand the entire system. For decision makers planning their own buildouts, that full-system discipline is no longer optional.
If you want to learn more about Nebius visit https://nebius.com/