Two racks of GPU accelerator cards, one in shadow with scattered dark idle cards and one lit with every card working, representing io.net keeping GPUs continuously in use.
Article Icon
Allyson Klein
@
TechArena
Oct 8, 2026

How io.net Rewrites the Economics of AI Compute for Startups

AI infrastructure contracts with hyperscalers can come with terms that are hard for smaller AI startups to navigate. On a recent podcast, I sat down with Ilkhom Sidikov, who works on the neocloud io.net’s IO Intelligence platform, to talk about how decentralized GPU networks are competing with hyperscalers by offering flexible, affordable options to AI innovators.

From Web3 Roots to an OpenRouter Listing

Ilkhom joined io.net in early 2024, not long after the company began turning a decentralized GPU network into something usable for AI startups and developers, an effort Ilkhom said got its start “in the age of Web3 and AI.”

He first worked on network infrastructure, and since 2025, his main project has been IO Intelligence, the vertical where io.net builds applications for AI startups, including model training, model inference, and video and image generation.

One recent milestone was being listed on OpenRouter, the LLM gateway many developers know. Io.net’s GPU network now hosts open-source models handling 2 billion to 5 billion tokens a day for OpenRouter.

Where Neoclouds Offer More Flexibility

Ilkhom said hyperscalers typically lock customers into one- to three-year contracts calibrated to protect their own margins, agreements that can be “really hard” on engineers and small AI startups. Under such contracts, a customer can end up paying for GPUs, like older 4090s, that fall behind newer architectures before the term ends. Even the AI grants that hyperscalers sometimes offer, he added, come with a catch. They typically require running on that provider’s own infrastructure.

Neoclouds, Ilkhom said, offer more room to negotiate, and io.net itself supports pay-as-you-go pricing for both GPU rental and LLM inference. Io.net’s specific advantage is how it negotiates use periods. Customers can rent a cluster for as little as 10 or 12 hours, and once they’re done, io.net shifts those same GPUs into its LLM inference infrastructure rather than leaving them idle.

“We never lose money on the idle GPUs,” Ilkhom said, adding that an idle GPU costs the same to run as one fully in use. That constant reallocation is what lets io.net offer clusters at a fraction of hyperscaler pricing.

Earning Enterprise Trust

Enterprise teams bring healthy skepticism to decentralized compute, particularly around reliability and security. Ilkhom said io.net encountered fraudsters early on who would alter a GPU’s metadata to make a cheaper card, like a 4090, pass as a far pricier H100.

Io.net’s fix was building software that profiles every hosted GPU on an hourly basis, checking whether each card meets performance benchmarks for what it claims to be.

On the network side, io.net connects nodes through virtual private networks (VPNs) to secure data in transit. It recently added a confidential inference feature that lets enterprise clients verify that every token generated, from the GPU level through the gateway to the back-end application, happens inside a secured environment.

Engineering for Scale and Profit

Running inference profitably at scale comes down to a handful of techniques, Ilkhom said: quantizing models down to formats like FP8 or INT4, using key-value (KV) caching to speed up throughput, and batching inference requests. Skip those optimizations, he said, and providers selling API tokens can end up burning cash, since deploying a raw model with default settings can make it hard to even break even.

Agents Raise New Demands

The rise of AI agents presents new demands on compute and data. Context windows have grown enormous, Ilkhom said, citing Anthropic’s own claim that nothing stops the company from building a model with a million-token context size. Combined with the development of MCP servers and skills that let agents pull real-time data instead of relying on static databases or vector search, agents now behave less like single inference calls and more like systems that maintain memory and reach out for information as they need it.

The TechArena Take

My conversation with Ilkhom conversation showed how decentralized providers like io.net are winning through an intensive focus on optimizing their resources. Io.net does this by keeping GPUs constantly busy rather than allowing them to sit idle between contracts.

Every technique he described, from shifting idle clusters into LLM inference to quantizing models, squeezes more usable compute out of the same hardware hyperscalers rely on too. Hourly hardware benchmarking helps build the trust enterprise customers need before relying on that hardware at all.

Ilkhom’s 5-year outlook does not predict that decentralized compute will displace the hyperscalers. After all, regulated and well-capitalized buyers have reasons beyond price to stay where they are. What his outlook does suggest is a market potentially splitting by workload, with contract-scale commitments on one side and short-duration, price-driven capacity on the other. For technology leaders, this opens up new possibilities to start matching each workload to the economics that fit it.

To learn more, visit io.net.

‍

Subscribe to Our Newsletter

Read the latest in the world of AI, data center, and edge innovation.