
The same chip architecture that Axelera AI built for the power, cost and latency constraints of edge vision AI can now also handle multiuser generative AI workloads at the edge, using the same Metis and Europa chip family and toolchain.
We sat down with Alexis Crowell of Axelera AI, which builds AI chips for edge devices. Alexis said that expanding capability matters as more AI workloads move beyond the data center and edge devices are asked to run increasingly complex models. Europa, she said, is built to run vision-language models and agentic AI directly on edge hardware.
We also talked about what digital in-memory computing changes about chip performance; how customers worldwide are using Axelera AI’s technology; and what it takes to offer a predictable cost per token as inference workloads grow.
Here’s what we learned.
A: Most of the industry has tried to scale AI from the data center out to the edge, and that direction has never really worked. Data center architectures are built for abundant power, abundant budget and abundant time. None of that is available at the edge. So the chips and systems built for one environment were getting forced into a location they were never designed for, which meant serious tradeoffs had to be made.
We believed there was a different way and started by holding those “limits” as real limits. We built an architecture that solves for small power envelopes, limited budgets and time from the beginning, because at the edge, those constraints are real.
Our first architecture, the Metis AIPU, had to hit real performance within a tight power envelope and a real price point, and that discipline is exactly what let us scale to the Europa AIPU without compromising the efficiency we perfected. We have even taken that edge architecture with both Metis and Europa and built server-class products supporting customers globally with full-length, full-height PCIe cards built around multiple AIPUs.
A: Customers are putting Metis to work solving real problems in places most people never think to look.
A large global convenience store chain uses it for real-time shelf monitoring, catching stockouts as they happen. An agritech company in New Zealand built it into automated apple sorters, and a healthcare technology company in Europe uses it inside an automated pill sorting machine, both cases where accuracy and speed directly affect people’s safety.
The utilization keeps expanding. A drone company in Croatia deploys Metis in a drone pack built for search and rescue missions. In India, a company built a cargo monitoring system for the country’s busiest port, bringing computer vision to logistics at scale.
These are just a handful of the deployments in production, with new uses for vision and localized models emerging around the world. What stands out is how different these use cases are from one another, and how each one found its way to edge AI because it was the only architecture that could meet their constraints on power, cost and latency at once. It’s a reminder of how much innovation is happening at the edge, in every corner of the world.
A: Traditional processors spend most of their energy and time moving data back and forth between memory and compute, especially for the matrix-vector multiplication that makes up roughly 80% of AI inference workloads.
Digital in-memory computing performs that math directly inside the memory cell, so the data barely has to move at all. That single change removes the biggest bottleneck in conventional architectures and turns into real gains in both speed and power. The RISC-V controlled dataflow is what makes this programmable and precise rather than a fixed-function shortcut. It gives us four independently programmable cores that can run parallel model execution with deterministic, digital accuracy, unlike analog in-memory approaches that trade precision for efficiency.
That combination is why the Metis architecture delivers up to 15 TOPS per watt and three times better performance per watt than GPU-based solutions, all while keeping accuracy indistinguishable from a full floating-point model. The result is an architecture that scales cleanly. The same principles that made the Metis AIPU efficient at 214 TOPS carry through to the Europa AIPU at 629 TOPS, which means moving the math into memory isn’t just a one-time efficiency trick. It’s the foundation of our roadmap.
A: The Europa AIPU is built on the same digital in-memory computing and RISC-V dataflow foundation as Metis, but with meaningful additions: twice the AI cores, native video decode, and 16 vector cores dedicated to pre- and post-processing, all of which make Europa perfectly suited for larger and more complex models and analytics. Those additions are what make native transformer support and multiuser generative AI workloads possible at the edge.
What stays constant for customers is the development experience. The same Voyager toolchain carries across both chips, so a team that built a pipeline on Metis isn’t starting over on Europa. That continuity is what turns an architectural leap into a growth path instead of a rebuild.
Metis remains an efficient powerhouse for computer vision use cases at the edge, including embedded systems where power and space are tightly constrained. Europa extends that foundation into a highly flexible AI accelerator built for a broader range of inferencing needs, from vision-language models and agentic AI to dozens of concurrent high-resolution video streams running natively at the edge. For customers, that means the same tooling and philosophy they already trust, now available across an even wider set of use cases.
A: Digital in-memory computing removes the data movement bottleneck that makes costs scale unpredictably as workloads grow, so the economics hold steady whether an operator is running a handful of streams or scaling to dozens of concurrent workloads.
The other separator is architectural flexibility. Operators locked into separate chips for CNNs and separate chips for transformer-based models face a cost structure that multiplies every time they add a new model type. A unified architecture that handles both on the same silicon keeps cost per token stable as workloads evolve from vision to generative AI, which is the discipline that will define the platforms operators can build a business on two years from now. Europa delivers that flexibility to customers.
Lastly, when models are run on one’s own infrastructure, the cost per token is much more manageable than running models in the clouds through providers that change cost structures at their discretion. With the Axelera Voyager toolchain, customers can use open-source models tuned to their weights and deploy them in-house. This includes Voyager Wingman, an agentic platform that can build pipelines based on natural language prompts.