
Customers are signing five-year contracts for AI silicon that has barely started shipping, while the models running on that hardware turn over in weeks, leaders from Crusoe, Microsoft and VAST Data said during a panel at AI Infra Summit 2026.
I had the delightful opportunity to moderate “Under the Hood of AI Infrastructure: What Top Leaders Do Differently,” with Erwan Menard, Crusoe’s senior vice president of product management; Donald Thompson, a distinguished engineer at Microsoft; and Alon Horev, co-founder and chief technology officer at VAST Data.
Across the hour, the three described an industry committing capital years ahead of knowing what it will run. Hardware gets committed in multiyear blocks and new power takes years to bring online, while the models and workloads running on them can change in weeks.
Erwan described clocks that do not agree: “I think in days. The data center guys used to think in years. Now, at places like Crusoe, they think in quarters. The energy world thinks in multiple years.”
One of Alon’s customers moves faster still, retiring the model its users are talking to every 30 minutes and publishing a newer one so that results stay fresh.
Alon called AI the “reset event of infrastructure globally,” and the technology hasn’t stopped resetting. Four themes from our discussion show where: Training and inference have folded into one continuous loop, the pace of innovation has turned into a design constraint, the megawatt has become the unit to plan around, and the data center is becoming a single computer spread across a building.
Inference is probably bigger than training across the worldwide GPU fleet, Erwan said, and at the AI-native companies Crusoe works with, it “becomes the lion’s share” once a product reaches volume.
Open models drive a lot of inference traffic too, he said. Crusoe helps customers modify a model with their own data, put it into production and retire it on their own schedule.
“The big motivation we hear from customers is control,” he said.
Donald carried the idea further. Digital workers are trained by people inside a specific work context, and their trajectories have to improve over time, which rules out pure training and pure inference.
“You have to be in a constant learning mode,” he said.
He described sleep-time compute, fine-tuning during quiet hours, and reinforcement learning in eval environments built to replay what an agent has just gone through.
Alon’s 30-minute customer is running that pattern in production. They kept pressing his team to speed up model loading, and he asked why they needed it hundreds or even thousands of times a day.
“Because we train nonstop,” they said.
Erwan told a story about an inference engine. Someone at the show that morning claimed bragging rights for building the first one three years ago, and Erwan doubted it was still in production.
“The pace at which these things are evolving is very annoying,” he said.
Crusoe shipped its first inference service a year ago and still tunes it weekly against fresh open source, new papers and new practices. Managed Kubernetes taught him the same lesson. A general-purpose version takes thousands of engineers and years; an AI-specific one has to be excellent at a short list of things, and Crusoe’s now powers 70% of its compute fleet.
Erwan described the new pace: Research reaches production in weeks, instead of two years, and his researchers talk directly to customers, who have seen Jensen Huang’s Pareto curve and are weighing throughput against latency.
“You need to walk with them and innovate with them on the fly,” he said.
Alon agreed that standards will likely hold you back at this pace, though he sees new ones forming organically, with SNIA and various companies working together. What changed, he said, is that buyers are less afraid of proprietary technology. Before AI, buyers wanted every layer of the stack to be generic, which made it hard to innovate or change big parts of it. “And AI changed that mindset at least for now.”
Most AI infrastructure gets bought in multiyear reservations now, Erwan said. He described five-year commitments for Vera Rubin, silicon that is only now starting to arrive. Buying that far ahead means committing to power before anyone knows what the workload will draw, and in reserved clusters already running, consumption sometimes lands well below what’s available. That gap opens room for a provider who understands both the machine’s power needs and the workload.
Closing the gap eventually means packing workloads onto shared hardware, in multitenant environments that absorb the peaks and valleys of inference traffic without stranding costly GPUs, he said. He predicted that by 2028, everyone will be optimizing for availability of megawatts.
“The energy wall is coming, and it’s going to take a lot of time to remove it,” he said.
Donald said Microsoft is working at the application and platform level to drive 24/7 inference, which is what pushes demand for megawatts.
Alon came at the same problem from the data side. He described the job of infrastructure builders as making “every megawatt, every terabyte on SSD and every CPU” serve training, fine-tuning and inference alike. He later asked whether reducing the doubled hardware some organizations carry on every stored byte could cut power along with it.
Packing workloads better only goes so far. Erwan expects power to stay short, since adjusting energy supply to demand takes years. Crusoe’s other answer runs through geography: hydro in Norway, geothermal in Iceland, and a small modular reactor paired with a modular data center that he hopes to put in service by mid-2027.
Alon called inference “a clustered problem,” with prefill and decode disaggregated so that “some GPUs do this; other GPUs do that.” Every open session ties up a block of GPU memory for as long as it lasts. That block is the KV cache. In some models, contexts of a million tokens push it to “tens or even hundreds of gigabytes,” crowding out other users, so it gets moved to a storage system or an SSD on another host. Routing then has to send that user back to whichever machine is holding their cache, so no single GPU completes the job alone.
“Everything is integrated into one logical computer that’s doing inference, … and KV cache is becoming a core component of it,” Alon said.
Eighteen to 24 months ago, Crusoe sized its data centers at four to five terabytes of storage per GPU, Erwan said. Today it is north of 10, and a given inference proof of concept takes roughly 20% fewer GPUs than a year ago.
He also wants the agent’s memory and context treated as its own layer of the stack. There is too much of it to hold in expensive memory, so some has to move to cheaper storage, and he sees “true infrastructural challenges in delivering that context.”
We should look at a data center now, Alon said at the end of the hour, as “a single computer just distributed across so many different machines and components.”
Hardware and power get committed for years. The models running on them change in weeks. That is the mismatch, and the panel treated it as a purchasing problem. Most of their answers were things to buy: longer silicon reservations, sites in Norway and Iceland, a reactor, more megawatts by 2028. One answer was operational, packing more work onto hardware already installed and already paid for, and it got less time.
That ordering is backward for anyone who has already signed. A five-year contract cannot be renegotiated into a different clock. What can still change is how the building runs: what gets packed where, what waits for a cheaper hour, what the machines actually draw against what the spec promised.
Donald raised this on stage, asking how intelligence gets applied to running the data center itself instead of leaving a person at a dashboard. Nobody answered. That is the question TechArena would put to vendors this year: What does the cluster do on its own when the model changes in six weeks? The answer may decide who is still competitive in 2028, when Erwan expects megawatt availability to be what everyone optimizes.