Latest AI news, expert analysis, bold opinions, and key trends — delivered to your inbox.
As demand for AI computing explodes, the biggest challenge is no longer training models—it's running them efficiently.
That's where French startup ZML is placing its bet. Backed by AI pioneer Yann LeCun, the company has launched ZML/LLMD, a free inference server that optimizes open-source LLMs across hardware from Nvidia, AMD, Google TPUs, Apple Metal, and Intel Arc. The goal is simple: help developers get the best possible performance from whatever AI chips they already have, instead of being locked into a single ecosystem.
Unlike ZML's earlier open-source framework, ZML/LLMD isn't open source. Instead, the company is offering it free to gather adoption and real-world feedback before deciding how to monetize it. Founder Steeve Morin says the priority is growth, not immediate revenue—a strategy made possible after raising $20 million from investors that include founders behind Docker and Hugging Face.
AI inference is quickly becoming the next battleground. As companies deploy AI at scale, reducing the cost and speed of running models can be just as valuable as building larger models. Tools that work across multiple chip vendors could weaken dependence on Nvidia and give developers more flexibility.
As AI infrastructure becomes just as important as AI models themselves, startups like ZML are proving that innovation doesn't have to come from Silicon Valley. If ZML/LLMD delivers on its promise, developers could soon have a faster and cheaper way to run AI—regardless of which chips power their systems.









