"Model to Metal"
Paris-based open-source project shipping a hardware-agnostic AI inference compiler stack that runs models on NVIDIA, AMD, TPU, and Trainium chips.
ZML is a Paris-based software company that builds an open-source, production-grade AI inference stack designed to decouple AI model deployment from proprietary hardware. Its core product, also called ZML, compiles machine learning models directly to multiple hardware accelerator targets, including NVIDIA GPUs, AMD GPUs, Google TPUs, AWS Trainium, and AWS Inferentia 2, allowing the same codebase to run across these platforms without requiring separate rewrites for each vendor's software stack. The project is built on the Zig programming language together with the OpenXLA and MLIR compiler infrastructure and the Bazel build system, and it is distributed under the Apache 2.0 open-source license through the public repository at github.com/zml/zml, which has accumulated over 3,500 GitHub stars and more than 700 commits. The company describes its engineering philosophy around three principles: explicit behavior over implicit behavior, composability over monolithic systems, and predictability over 'magic' abstractions, and it markets the product with the tagline 'Model to Metal' and the promise of running 'Any model. Any hardware. Zero compromise.' The toolkit supports popular open model families such as Llama and Qwen, and includes a virtual file system layer that can load model weights from local storage, Hugging Face repositories, HTTPS endpoints, or Amazon S3, alongside built-in sharding and distributed-execution support for scaling inference across multiple accelerators. In 2026 the company released ZML/LLMD, a further inference-serving tool aimed at speeding up AI inference workloads across heterogeneous chip types, which it distributes for free; it also shipped supporting tools such as zml-smi, a monitoring utility, and reported performance improvements including a claimed tenfold speedup in tokenization. ZML maintains a technical blog documenting these releases, operates a public Discord community for developers, and publishes documentation at docs.zml.ai. The company's GitHub organization profile lists its location as France and identifies engineers including a contributor using the handle r-chong among its visible team. ZML positions its target users as engineering teams that need to run large language model and other AI inference workloads in production without being locked into a single hardware vendor's software ecosystem, competing in the space of hardware-agnostic ML compiler and inference-serving tools alongside projects such as vLLM and TensorRT-LLM.
Primary type & automation activities this supplier delivers:
Production inference stack compiling AI models directly to NVIDIA, AMD, Google TPU and AWS Trainium from one codebase.
Free self-contained LLM inference server running Llama, Gemma, Qwen and Mistral across 5 chip architectures with DFlash decoding.
Company Facts
Founded
-
Primary Role
Software/Algorithm
Company Size
employees 50-100
Primary Region
Europe
Annual Sales
-
Funding Stage
-
Funding Total
-
ZML, a hot French AI startup endorsed by Turing Award winner Yann LeCun, has now released ZML/LLMD, software that could make running AI less costly.
TechCrunch Jul 08, 2026 AI Fundraising Startups AI inference Exclusive ZML