
Leading up to the AI Infra Summit in September, TechArena has spent the summer talking with the companies building AI infrastructure, from hyperscalers to storage, memory and networking, through to the power, cooling, and silicon that hold these systems together.
In our latest Five Fast Facts Q&A, we chatted with Tony Pialis, EVP and GM of Data Center at Qualcomm. He explains what changes when a company that grew up counting every milliwatt in mobile brings that discipline to inference. We discussed the memory wall, why data movement now costs more than compute, and what it takes to sell a rack instead of a chip. Here's what we learned.
A: Qualcomm’s perspective is different because we grew up solving compute problems under severe power constraints. In mobile, every milliwatt matters because battery life, thermals, and form factor are critical. That mindset translates directly to our AI inference solutions, since the real challenge in modern data centers is no longer peak compute, but delivering the most AI work within a constrained power and investment envelope.
Designing for efficiency first enables us to optimize the entire rack around cost-per-token and power efficiency rather than chasing benchmark peaks. Qualcomm Dragonfly combines specialized CPUs, AI accelerators, memory innovation through High Bandwidth Compute (HBC), and advanced connectivity in a disaggregated architecture designed specifically for inference. Our multi-generation roadmap is focused on maximizing performance per watt, improving token economics, and lowering total cost of ownership at scale.
A: The biggest constraint in AI inference isn’t raw compute anymore—it’s moving data efficiently. Model sizes are growing far faster than memory bandwidth and capacity, creating what the industry increasingly describes as the memory wall. Qualcomm believes that simply adding more compute does not solve the problem if memory becomes the bottleneck.
That realization led us to develop HBC, a near-memory computing architecture that brings compute and memory much closer together. Instead of shuttling massive amounts of data back and forth, HBC performs more processing near memory, reducing movement, lowering power consumption, improving effective memory bandwidth, and lowering overall system costs. HBC is poised to deliver significantly higher bandwidth-per-watt and capacity-per-watt compared with traditional approaches. The outcome is infrastructure designed specifically for modern AI inference workloads where efficiency, scalability, and predictable economics matter as much as performance.
A: I would characterize Dragonfly as a portfolio of rack-scale platforms rather than a single liquid-cooled product. Depending on the deployment, Dragonfly systems can support air or direct-liquid cooling. What is important is that customers increasingly evaluate AI infrastructure at the system level—not as isolated chips.
AI infrastructure is increasingly a systems problem rather than a chip problem. Customers no longer evaluate silicon in isolation—they care about rack-level performance, power consumption, networking, software orchestration, cooling, and overall cost of ownership. Qualcomm Dragonfly reflects that reality by bringing CPUs, AI accelerators, memory architecture, connectivity, software, and custom silicon together into a unified data center platform.
Optimizing one component is not enough. The infrastructure must work as an integrated system. One persistent challenge for operators is balancing compute, memory, networking, and power at scale. Bottlenecks often emerge not from lack of processing power but from inefficient data movement, infrastructure complexity, or underutilized resources. Our approach is to simplify this by delivering a rack-scale platform designed around inference efficiency, open software, disaggregated compute that help customers scale economically as agentic AI dramatically increases token demand.
A: To operate as a single platform, every layer must be designed around common system objectives: performance, efficiency, openness, and scale. Qualcomm’s strategy combines CPUs, AI accelerators, connectivity technologies, custom silicon, orchestration software, and developer tools into a unified architecture optimized for AI inference. The software layer is especially important because it coordinates workloads across heterogeneous compute resources and abstracts complexity from developers and operators.
At the same time, we believe openness matters. AI infrastructure is becoming increasingly heterogeneous, with CPUs, GPUs, XPUs, and specialized accelerators coexisting. The real challenge is enabling those components to work together efficiently. The seams that remain are often industry-wide—not unique to Qualcomm—including interoperability across different hardware ecosystems, software stacks, and deployment environments. Our goal is to minimize those seams through open standards and unified software rather than proprietary lock-in.
A: We believe the industry is undergoing a fundamental shift from measuring AI infrastructure by peak FLOPS to measuring it by tokens per watt and ultimately cost per token. As agentic AI drives massive growth in inference requests, economics become the defining factor. Operators need infrastructure that can deliver consistent throughput while managing power, cooling, and hardware utilization efficiently.
Memory architecture plays a central role because data movement increasingly consumes more energy than computation itself. That’s why Qualcomm invested in High Bandwidth Compute, which is designed to reduce energy consumed moving data, improve effective bandwidth, and lower total cost of ownership. But cost per token is ultimately determined at the rack level as well. Compute, memory, software orchestration, networking, and cooling must be optimized together. Qualcomm Dragonfly was built around that systems view, using a disaggregated rack-scale architecture to maximize utilization and efficiency. In our view, the winners in AI inference will be those who can consistently deliver the best performance-per-watt, performance-per-dollar, and long-term economics—not simply the highest headline specifications.