ZML

ZML

"Model to Metal"

Paris-based open-source project shipping a hardware-agnostic AI inference compiler stack that runs models on NVIDIA, AMD, TPU, and Trainium chips.

Visit Website

Company Overview

ZML is a Paris-based software company that builds an open-source, production-grade AI inference stack designed to decouple AI model deployment from proprietary hardware. Its core product, also called ZML, compiles machine learning models directly to multiple hardware accelerator targets, including NVIDIA GPUs, AMD GPUs, Google TPUs, AWS Trainium, and AWS Inferentia 2, allowing the same codebase to run across these platforms without requiring separate rewrites for each vendor's software stack. The project is built on the Zig programming language together with the OpenXLA and MLIR compiler infrastructure and the Bazel build system, and it is distributed under the Apache 2.0 open-source license through the public repository at github.com/zml/zml, which has accumulated over 3,500 GitHub stars and more than 700 commits. The company describes its engineering philosophy around three principles: explicit behavior over implicit behavior, composability over monolithic systems, and predictability over 'magic' abstractions, and it markets the product with the tagline 'Model to Metal' and the promise of running 'Any model. Any hardware. Zero compromise.' The toolkit supports popular open model families such as Llama and Qwen, and includes a virtual file system layer that can load model weights from local storage, Hugging Face repositories, HTTPS endpoints, or Amazon S3, alongside built-in sharding and distributed-execution support for scaling inference across multiple accelerators. In 2026 the company released ZML/LLMD, a further inference-serving tool aimed at speeding up AI inference workloads across heterogeneous chip types, which it distributes for free; it also shipped supporting tools such as zml-smi, a monitoring utility, and reported performance improvements including a claimed tenfold speedup in tokenization. ZML maintains a technical blog documenting these releases, operates a public Discord community for developers, and publishes documentation at docs.zml.ai. The company's GitHub organization profile lists its location as France and identifies engineers including a contributor using the handle r-chong among its visible team. ZML positions its target users as engineering teams that need to run large language model and other AI inference workloads in production without being locked into a single hardware vendor's software ecosystem, competing in the space of hardware-agnostic ML compiler and inference-serving tools alongside projects such as vLLM and TensorRT-LLM.

Capabilities & Activities

Primary type & automation activities this supplier delivers:

Applications & Industries

Product Categories

Products & Solutions

ZML
  • Launch Year: 2024

Production inference stack compiling AI models directly to NVIDIA, AMD, Google TPU and AWS Trainium from one codebase.

ZML/LLMD
  • Launch Year: 2026

Free self-contained LLM inference server running Llama, Gemma, Qwen and Mistral across 5 chip architectures with DFlash decoding.

Partnership & Notable clients

Associating Events

Contact ZML

WEBSITE

https://zml.ai

HEADQUARTERS

Paris , Île-de-France  75000

France

Company Facts

Founded

-

Primary Role

Software/Algorithm

Company Size

employees 50-100

Primary Region

Europe

Annual Sales

-

Funding Stage

-

Funding Total

-

Listed in RobotToday Supplier Discovery. Verified Profile

Related Coverage

Hot French startup ZML releases free product to speed inference across lots of AI chips

Hot French startup ZML releases free product to speed inference across lots of AI chips

ZML, a hot French AI startup endorsed by Turing Award winner Yann LeCun, has now released ZML/LLMD, software that could make running AI less costly.

AI Fundraising Startups AI inference Exclusive ZML