Tokens falling in a data center hall
Article Icon
Rachel Horton
@
TechArena
Sep 17, 2026

AI Infra Summit 2026: Why the Inference Buildout Isn’t Pausing

AI infrastructure leaders from around the globe gathered this week for the 2026 AI Infra Summit in Santa Clara against a backdrop of AI model slowdown discussions. Speakers from Crusoe to OpenAI separated the two issues: the slowdown call concerns frontier model training, and the infrastructure they are building runs inference.

Across three days at the summit, operators and suppliers pointed to the continued growth of inference demand, stating that they can build only as fast as they can secure power. They laid out how to get more tokens from each megawatt through rack-level power management, campuses that ramp with the grid, and cooling that uses no water.

The Slowdown Question

Ed Nelson, co-founder of the summit, put the question to Lip-Bu Tan, CEO of Intel, on the main stage: does the slowdown call mean the infrastructure buildout has to pause, or does it shift the focus toward inference at scale? Tan responded by saying he’s seen multiple cycles of technology growth, but this one is bigger.

“…Some correction is kind of natural, and I don't overreact to it,” Tan said. “But something that you have to pay attention (to): We live in the Silicon Valley. It’s a bubble. And we all get so excited about AI. When you go to a different part of the country…you find that there's a very strong anti-AI sentiment.”

Tan said leaders in the industry have to address the negative side directly while continuing to adopt the technology, prioritizing security and ensuring that data and privacy are not compromised.

The Workload Changed

Ian Buck, VP of hyperscale and HPC computing at NVIDIA, put numbers on how much the inference workload has changed. The 2023 chat workload that became the industry benchmark ran about 1,000 input tokens, a 4,000-token cache and about three turns, with a person reading each answer before asking the next question. The agentic benchmark Buck cited averages 142,000 input tokens. He described agents running 65 turns to reach an answer. Commercial services already run thousands of sub-agents, he said. He called the new workload 100 times more demanding.

“Agents have now taken humans out of the loop,” Buck said, “and the compute will be consumed as fast as the agent can make the turn.”

Demand is not slowing

Dave Patterson, distinguished engineer and Google Fellow, described hardware staying in service long past its amortization schedule inside Google. TPU v2 and v3, chips from eight and ten years ago, still run inference. Google amortizes accelerators over six years and CPUs over eight, and Patterson said the hardware will last longer than that.

“There’s just such a desperate need for cycles that people are absolutely keeping them beyond the amortization,” Patterson said.

Bloomberg’s Dina Bass asked Sachin Katti, VP of compute strategy and GPT-Infra at OpenAI, whether the safety and alignment discussion affects OpenAI’s compute plans. He said it increases the need for compute. Making very capable models safe requires safety models and alignment models. Those have to be trained and served as well.

“If anything, I think it actually double underscores the need for more compute,” Katti said.

During a TechArena panel, an audience member asked when the industry gets its next moment to pause and reset expectations. Erwan Menard, SVP of product management at Crusoe, compared the current supply imbalance to the flooding of hard disk drive factories in Thailand after a monsoon in 2011; at the time, Thailand manufactured ~40% of the world's hard disk drives. He called it the closest precedent in his more than 25 years in infrastructure.

“What we’re experiencing now is way bigger than that at every layer of it, which tells us that demand is essentially insatiable,” he said.

Two years ago, Menard said, the question at a session like this would have been what a provider planned to do with the Hopper GPUs on its balance sheet. NVIDIA was shipping a new generation every year, and there was no time to monetize the old silicon before the amortization schedule ran out. Two generations later, the GPU-hour price of a Hopper cluster has risen since the beginning of the year. Menard called that another validation of demand.

Donald Thompson, distinguished engineer at Microsoft, took the audience question after Menard.

“A lot of it is irrational,” Thompson said. “And I think it’ll take an exogenous event to create a pause. I don’t think that the market itself will pause. Even if it’s in our best interest to pause, I don’t think we will self-regulate in that way.”

Asked what kind of event, he said geopolitical.

Inference is Production

John Roese, global CTO and chief AI officer at Dell, described what it took to get AI into production inside his own company. Dell had 900 AI projects when he took the role. The company discarded them and restarted with four strategic projects covering sales, services, supply chain and engineering. Roese put the result at $30 billion in revenue growth over two years alongside a decline of about 7% in absolute cost.

“To get an enterprise into production is non-trivial,” he said. “Many of the things you have to deal with have nothing to do with the technology.”

Power is the Constraint

NVIDIA’s Ian Buck said the company does not think in terms of absolute token performance.

“The metric that matters is your total token throughput per megawatt,” Buck said.

Buck described software that manages the power of each rack in real time and coordinates with the data center management system. On one model he cited, GPU power ranged from 600 watts to 1,900 watts. Operators provision racks for that peak. Buck said the gap between provisioned power and actual draw is capacity operators are not using. Managing to the actual draw lets them install more racks inside the same provisioned power. He put the gain at up to 40% more compute in the same megawatts.

The Energy Wall

Menard described the constraint from the operator side. Crusoe sources energy, builds data centers, manages GPUs and serves tokens. He said the company optimizes for the value it extracts from each megawatt.

“If you look at 2028, what you’re optimizing is availability of megawatts,” he said. “So, you need to be ready for that because the energy wall is coming and it’s gonna take a lot of time to remove it.”

Crusoe may offer its customers infrequent service degradations to avoid stressing the grid, Menard said, because the company shares that energy with the broader community.

Katti described a similar approach to the grid at the 8-gigawatt Ohio campus OpenAI is building with NVIDIA. Jobs can ramp down when grid demand spikes and ramp up when it drops.

“This can help stabilize the grid in Ohio,” Katti said.

Buck said a scheduler that knows where each job runs and how much power it draws can balance delivery across rows of the Ohio site. He put the gain at 30% to 40% more GPUs inside the same allocation. The site’s power system was designed for that control, from the compute trays up to the transformer.  

More From What Is Already Built

Katti also described a second source of efficiency: a model tuning its own inference. OpenAI took a checkpoint of its Astra model, already optimized for Blackwell, and ran it on Rubin with minor changes. Katti said it ran three times faster out of the box. The team then had the model optimize its own inference serving on Rubin. Over 72 hours, Katti said, the model produced another 2x. He said that work would have taken weeks or months in the past.

Two announcements applied the premise of more tokens per megawatt to installed infrastructure. Axiado previewed a platform efficiency controller that runs AI agents on management silicon. The agents scale processor voltage and frequency and drive cooling from measured conditions rather than worst-case assumptions. Hammerhead AI announced a distribution agreement with TD SYNNEX for a 1-megawatt SKU of its power orchestration software. The offering targets operators facing utility interconnection queues that run several years in major U.S. markets.

Nick Harris, founder and CEO of Lightmatter, described a different route to more capacity without new power. Optical interconnect spans a kilometer, so operators can spread current-generation GPUs across existing facilities rather than concentrating them in new high-density racks.

“It’s reusing data centers that are already there, because optics doesn’t care,” Harris said. “A kilometer reach is not a problem.”

The Water Question

Robert Hormuth, corporate vice president of architecture and strategy at AMD, addressed the public reaction to data center siting, power usage and water usage. He said the industry needs to do a better job governing all three.

“I don’t know if I want one in my backyard either,” Hormuth said.

Airsys, a global provider of specialized thermal management and engineered cooling solutions, won TechArena’s inaugural Ad Astra: AI Infrastructure Competition for LiquidRack ahead of the summit. LiquidRack integrates cooling into the rack rather than relying on facility infrastructure. The system uses patented spray cooling, which Airsys says transfers heat up to three times more efficiently than conventional methods. For data centers running LiquidRack, Airsys reports a PUE below 1.02, no water consumption and 80% less dielectric fluid than immersion cooling. Jeff Moore, VP of strategic partnerships at Aegis Cooling, an Airsys company, joined the Day 1 panel on power and cooling at scale.

The TechArena Take

The slowdown conversation and the infrastructure conversation ran on separate tracks in Santa Clara, and operators should keep them separate. Frontier labs pacing capability gains does not reduce the compute required to serve models already in production. The safety work those labs described adds to it.

What changed at this summit is the unit of account. Tokens per megawatt showed up in NVIDIA’s keynote, in Crusoe’s operating model and in the claims vendors brought to the floor. Operators plan around the power they can get. Rack-level power management, grid-aware scheduling, models tuning their own inference and cooling that uses no water all add capacity without adding megawatts.

Crusoe's Erwan Menard gave the timeline worth watching. Software moves in days, he said, and the energy world moves in multiple years. Operators planning for 2028 are not optimizing for chips. They are optimizing for the megawatts they can secure.

Subscribe to Our Newsletter

Read the latest in the world of AI, data center, and edge innovation.