Row of server motherboards with glowing blue and magenta circuit traces.
Article Icon
Chloe Jian Ma
@
Epic Semi
Sep 29, 2026

The Agentic AI Bottleneck Keeps Moving

Why enterprise infrastructure must optimize the entire workflow, not only model inference

A recent study of five representative agentic AI workloads found that the dominant bottleneck shifted from retrieval to code execution, scientific tools, summarization, and model inference. In several tests, work outside the model consumed more than 80% of end-to-end latency.

That matters because an AI agent does much more than running a model. It plans, retrieves data, calls tools, updates its state and repeats the process until the task is complete. Each stage stresses a different part of the system.

For banks, hospitals, universities and national labs that buy and run their own infrastructure to keep data on premises, optimizing only the accelerator may only yield very limited impact on the entire workflow, or even worse, leave your most expensive and power-hungry accelerator waiting.

The bottleneck moves

The study, “Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective,” published in November 2025, instrumented the five workloads end-to-end across two systemsi. The results do not converge. In each workload, a different stage consumed the largest share of end-to-end latency:

Selected results from two systems and multiple benchmarks. The dominant stage changes with workload inputs and hardware configuration.

  • Only in the last case is the bottleneck the matrix math the industry has been building for.

For enterprise infrastructure teams, that ratio changes the buying decision. The bottleneck moves with the work, and a production deployment runs several kinds of work at once. No single resource can be scaled up to fit the whole fleet. Take the retrieval-augmented pipeline from the study, where nearest-neighbor search accounted for more than 80% of the time and runs on CPU cores. Buying more accelerators speeds up the model inference in the remaining share. The search step, which decides how long the whole job takes, does not get faster at all.

One agent needs several kinds of compute

Look one layer down and the same pattern holds. At the model level, prompt prefill benefits from high parallel compute throughput, while token-by-token decode depends more heavily on memory bandwidth and capacity. Outside the model, data processing, tool calls, orchestration and security run efficiently on general-purpose cores. An agent moves between these phases continuously.

Balanced systems are becoming a mainstream design priority. Intel CEO Lip-Bu Tan told the Computex 2026 audience that agentic AI is “returning the CPU to a position of prominence,”ii and AMD made the same case for CPUs in the AI data center earlier in the year.iii The single-silicon era of AI infrastructure is closing.

Why we built Contrail AIX

Organizations and nations that desire to gain greater control over how AI is deployed, customized and governed keep their data on premises, deploy open-weight models and build infrastructure for the complete agent workflow deployable inside their own walls. That is what Epic Semi designed Contrail AIX to be.

What we built to run AI agents

Contrail Compute, the RISC-V CPU and AI superchip inside our Contrail AIX platform, puts 32 high-performance RISC-V cores and 16 RISC-V AI acceleration cores on one die, sharing one memory and storage space across a common on-chip fabric.

This architecture keeps retrieval, orchestration, model execution, and tool calls close to the same agent state, reducing unnecessary data movement between separate devices. Each Contrail Compute chip can serve as a native agentic AI execution unit, and these units are small, complete, and fungible. Small means granular enough to scale independently. Complete means useful work can run locally. Fungible means any unit can be scheduled, replicated, drained, or replaced through a common operating model.

Agent-aware orchestration compiles the workflow into a physical plan, places each fragment on the right resources, preserves data locality, and continuously resizes the fleet. Units pool when demand rises, fail over automatically, and disappear when no longer needed.

We built it on RISC-V so a customer’s software investment outlives any single vendor’s roadmap and any single government’s export policy. Contrail AIX carries a 5A992C export classification, which reaches substantially all markets. It supports the RISC-V hypervisor extension with full hardware virtualization and IOMMU, RVA23-compatible architecture, trusted secure boot and full-chip security management. Ubuntu came up shortly after first silicon, and llama.cpp and open-weight models are running today. Compute is the first product on a multi-generation roadmap, from a team that has ramped millions of microprocessors.

Agentic benchmarking is underway. We will publish performance results after optimization and validation so customers can evaluate defensible end-to-end measurements rather than synthetic claims.

If you operate AI on premises and want control of your data, bring us your workflow bottlenecks. We’ll help you measure them end-to-end on Contrail AIX.
‍

[1] Ritik Raj, Souvik Kundu, Ishita Vohra, Hong Wang and Tushar Krishna, "Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective," arXiv:2511.00739, submitted Nov. 1, 2025, revised April 16, 2026, https://arxiv.org/abs/2511.00739.

[1] Intel Newsroom, "Computex 2026: An Intelligent World Built on Silicon," June 2026, https://newsroom.intel.com/artificial-intelligence/computex-2026-an-intelligent-world-built-on-silicon. Quote as reported in Divyanshi Sharma, "Computex 2026: Intel announces Xeon 6+ processors, says AI will make CPUs important again," Digit, June 2, 2026,

[1] AMD, "Agentic AI Brings New Attention to CPUs in the AI Data Center," AMD corporate blog, March 13, 2026, https://amd.com/en/blogs/2026/agentic-ai-brings-new-attention-to-cpus-in-the-ai-data.html.

‍

Subscribe to Our Newsletter

Read the latest in the world of AI, data center, and edge innovation.