
By Rashid Attar, SVP, Engineering, Qualcomm Technologies, Inc.
Originally published July 30, 2026, on the Qualcomm OnQ Blog (10 min read)
Part one of this blog post compared three AI-inference memory strategies — SRAM, HBM and HBC. Since decode is memory-bound, HBC wins by computing near data, relaxing interface bandwidth, capacity, power and cost trade-offs instead of pitting them against each other. Part 2 presents an architect’s view and an honest comparison for a full decode solution.
The language computer architects use to reason about performance is the roofline. The model plots attainable performance against a workload’s arithmetic intensity — the number of operations performed for every byte moved. A workload can only reach peak compute once it does enough arithmetic per byte to hide the cost of fetching that byte. Below that point it sits on a sloped roof whose height is set entirely by bandwidth.
This is exactly where autoregressive generation lives. Producing each token reads a large volume of parameters and context but does only a small multiply-accumulate with them — an arithmetic intensity of roughly one to two operations per byte at low batch sizes. Batching lifts it somewhat, but latency targets and the memory cost of holding more context cap how far. The workload is pinned to the bandwidth roof — and no amount of extra arithmetic silicon moves it.
Here is the architect’s insight. Arithmetic intensity is not a fixed property of the model; it depends on which boundary you measure it across. The operation count is the same for everyone, so the way to raise intensity is to shrink the denominator — the bytes that must cross the expensive boundary between memory and compute. That is precisely where the three approaches differ.
Seen this way, the debate stops being about peak throughput and becomes a single architectural question: which design raises the arithmetic intensity at the boundary that limits you? The roofline says that is the only quantity that moves memory-bound performance — and HBC is the one approach that raises it structurally, rather than renting more bandwidth to feed the same wasteful crossing.
Previous blog posts in this HBC series argued that AI inference is limited by how well a system feeds its compute, and that HBC relaxes the trade-off between bandwidth, capacity, power and cost. That argument is sound — but told at the level of a whole system, it can flatten an important truth: no single memory approach wins everywhere.
Modern inference is not one workload. It is a mix of very different kinds of work, and the best architecture for one is not the best for another. A fair comparison has to say where each approach genuinely leads, ties or trails — and only then draw the system-level conclusion. That is the point of this piece.
Generating a token exercises two dominant kinds of computation, and they stress memory in opposite ways.
Add batch size as a second axis. At low batch and tight latency — a single user waiting on a fast response — there is very little arithmetic to amortize each byte fetched, so the work is at its most memory-bound. As batch size grows, more useful work rides on each fetch, and the picture shifts again. Any claim about who “wins” is meaningless until it names the workload and the batch regime.
With those regimes named, the comparison becomes concrete. The table below is deliberately even-handed — it hands each approach the wins it has earned.
Keeping the model in very fast on-die memory delivers extraordinary bandwidth at very low latency. In the regime it is built for — FFN layers at low batch size and tight latency — this is genuinely the approach to beat. It is typically the least costly per unit of throughput where it applies, it matches or exceeds HBC on energy per token, and it clears HBM-based accelerators comfortably on both.
The limits are structural, not incidental. Sustaining that on-chip bandwidth imposes strong constraints on the compute design, and the memory works well only for static, fixed-size data — which is exactly why it excels at FFN work and does not extend to attention's growing, irregular footprint. Capacity per device is small, so a large model must be spread across many devices, carrying a large threshold investment and real system complexity. It can be thought of as a narrow specialist.
Surrounding a large processor with stacked HBM buys both generous capacity and good bandwidth, which is why this approach runs a broad range of workloads today. It handles FFN and attention networks, short context and long, without falling off a cliff. Its ceiling is the boundary itself: every byte the computation touches must still be read out of the memory and driven across the interface, and that bandwidth is bought through expensive, supply-constrained packaging. It is rarely the outright winner on any one regime, but it is also seldom the big loser — the safe generalist.
Computing the most data-intensive work inside the memory changes what has to move. On attention — whose ever-growing context is precisely the traffic that dominates long-running chats and agents — doing the work in place, and sending out only compact results, is decisive. Avoiding data movement rather than merely managing it compounds the advantage at system scale, where data movement is the dominant energy cost and the binding physical limit.
Where it does not claim an outright win is the specialist's narrow home turf: on low-batch FFN work, a well-matched on-chip design can equal or better it. HBC is competitive, but not the peak.
If each approach excels in a different regime, how can there be a system-level winner at all? Because a real deployment does not run one regime — it runs a weighted mix, and the weights are moving.
As models take on longer context and more agentic, multi-turn use, attention's share of time and energy rises. The regime where on-chip memory is strongest — small, fixed, low-batch FFN — is a shrinking slice of the total; the regime where HBC is strongest is a growing one. The weighted average tilts accordingly.
So the fair statement is a careful one. On tokens per watt and tokens per dollar, the specialist can win an individual layer, but the architecture that minimizes data movement wins the realistic mix — and wins it by more as context grows. The claim is a weighted one, not a clean sweep, and it is stronger for being stated that way.
As AI inference workloads evolve, no single memory architecture wins every scenario. When HBC, HBM and on-chip SRAM are viewed through the lens of the roofline model, each excels in different workload regimes. While on-chip SRAM leads in low-batch FFN processing and HBM remains a versatile generalist, HBC’s ability to perform computation where data resides fundamentally reduces data movement, making it the strongest approach for attention-heavy workloads and increasingly favorable as context lengths continue to grow.
Inference is a mix, and the mix is shifting toward exactly the work that rewards minimizing data movement. That is why the system-level conclusion survives honest accounting: not because one approach wins every regime, but because HBC wins the regimes that increasingly matter most — and wins them by more over time.
The future economics of AI infrastructure will therefore be determined not by peak compute alone, but by which architecture does the most useful work per byte moved. The next blog post in this HBC series turns from why the advantage exists to what it takes to build it at scale.
No, not for a deployment that covers a broad set of AI Inference applications. On-chip SRAM wins its regime — low-batch FFN layers at tight latency — but that regime is a shrinking slice of the total workload mix as context lengthens and agentic, multi-turn use grows. The architecture that minimizes data movement wins the weighted average, and it wins by more as that average shifts toward attention's growing, irregular footprint.
Read the first blog post of the series for an introduction to HBC
In an HBM-based accelerator design, more bandwidth still requires every parameter byte to cross the interface each time it is used — it rents more of the same expensive crossing rather than eliminating it. HBC does the bandwidth-heavy work where the data already sits, so what crosses the boundary is only compact results; the denominator shrinks, the arithmetic intensity at the boundary rises, and that is the only quantity the roofline rewards for memory-bound workloads. It is a structural change to what moves, not a larger pipe for the same movement.
Learn more about Qualcomm's AI accelerators using Qualcomm HBC
The system-level implication is already in the weighted verdict: data movement is the dominant energy cost and the binding physical limit at scale, so the architecture that reduces it structurally — rather than managing it more efficiently — holds the compounding advantage as inference workloads grow longer and more complex.
Explore the Qualcomm Dragonfly data center portfolio