
Every new AI data center hits the same wall before it runs out of ambition: power. Operators can order more accelerators, but the grid connection, the cooling budget and the rack density are fixed. That ceiling is quietly rewriting what gets built, and it has turned efficiency from a line on the TCO sheet into the thing that decides how much useful AI work a site can deliver.
As part of our summer series with the companies building the AI stack ahead of the AI Infra Summit this September, I sat down with Eddie Ramirez, VP of marketing for Arm's Infrastructure Business. Arm has spent years making the case for performance per watt in the data center, and that case lands differently now that power is the hard limit rather than a background concern.
We talked about why the CPU is becoming the control plane for an entire rack, what open chiplet standards like UCIe and CHI have to hold together as AI silicon turns into an assembly of parts, and how operators can keep pulling useful work out of a data center over a lifespan that will outlast several generations of models. Here's what I learned.
A: AI infrastructure is increasingly constrained by fixed power, cooling and rack density rather than demand for compute. As AI deployments scale, operators are no longer optimizing individual servers. They're optimizing entire racks, and ultimately entire data centers. Efficiency therefore isn't simply about lowering TCO, but it is about how much useful AI work can be delivered within a fixed power envelope. As AI workloads become more continuous and agentic, maximizing performance at the rack level, not just within a single server, will increasingly define competitive AI infrastructure.
A: AI systems are evolving from executing individual models to coordinating fleets of specialized models that are continuously reasoning, retrieving information, calling tools, managing memory and moving data across large clusters of compute pools. Those continuous system-level tasks naturally belong on the CPU. As AI infrastructure scales, the CPU becomes the control plane for the entire rack, coordinating data movement, scheduling work, feeding accelerators efficiently, and ensuring entire system resources are fully utilized. The result isn't a competition between CPUs and GPUs, but a better-balanced system where each processor is optimized for the work it does best – and work is assigned to the best suited processor.
A: Open standards like UCIe and AMBA CHI are essential because they allow compute, AI accelerators, memory, and networking to evolve independently while still working together as a cohesive system. That gives silicon providers the flexibility to innovate faster without sacrificing software portability or ecosystem compatibility. That's also the thinking behind Arm's Foundation Chiplet System Architecture (FCSA), which we contributed to the Open Compute Project to help establish an open, interoperable foundation for next generation chiplet-based AI infrastructure.
A: Building AI silicon at the leading edge is extraordinarily complex – beyond the digital design, advanced node processes and packaging layer on additional challenges. Everything we can do to simplify our partner’s journey, accelerate the path to silicon and de-risk production holds a great deal of value. With Arm Compute Subsystems, partners receive a production-ready compute foundation that's already validated for performance, software compatibility, and system integration. That allows engineering teams to focus their investment on the capabilities that differentiate them, whether that's AI acceleration, networking, memory architecture, or custom system innovation. The result is faster time to market, lower development risk, and broader innovation across the AI ecosystem.
A: AI models will evolve far faster than the infrastructure that supports them. Data centers are long-term investments that are both capital and labor intensive, so the compute layer must continue delivering value as workloads, models, and software stacks change over time. That means balancing performance, energy efficiency, memory bandwidth, software portability, and ecosystem support at rack scale, not simply optimizing for a single benchmark or today's leading model. The most successful AI platforms will be those that consistently deliver more useful AI work within fixed power and space constraints as AI continues to evolve.