"Own Your Specialized Intelligence"
Redwood City AI infra startup (2022) serving open-source LLMs like Kimi K3 via API, used by Cursor, Vercel and Notion.
Fireworks AI is a generative-AI infrastructure company founded in 2022 and headquartered in Redwood City, California. It was started by a team of seven engineers with backgrounds building and maintaining PyTorch and large-scale machine learning infrastructure at Meta and Google, including CEO Lin Qiao, formerly head of PyTorch at Meta. The company operates a training and inference platform that lets organizations deploy, fine-tune and serve open-source large language models -- including DeepSeek, GLM, Kimi, Qwen and Minimax -- through OpenAI- and Anthropic-compatible APIs. Its inference offering spans serverless pay-per-token access (priced from roughly $0.07 to $3 per million tokens depending on model and priority tier), on-demand dedicated deployments, and reserved capacity for high-volume workloads, running across multiple cloud regions on NVIDIA and AMD hardware. On the training side, Fireworks AI supports LoRA and full fine-tuning, guided training runs with cost estimation, and integrated reinforcement-learning pipelines for building custom or task-specific models. In August 2026, Fireworks AI began hosting Moonshot AI's Kimi K3 model on Microsoft Foundry, extending its role as a third-party inference backend embedded in a major cloud AI marketplace. Companies publicly named as customers on the Fireworks AI website include Cursor, Vercel, Notion, UiPath, Sourcegraph, Genspark and Motif. The company has raised venture funding from Benchmark, Sequoia Capital, Lightspeed Venture Partners and Index Ventures, with NVIDIA and AMD also participating as strategic infrastructure investors, and has been reported at a multi-billion-dollar valuation following funding rounds through 2025. According to a 2026 employer-profile snapshot, Fireworks AI has approximately 63 employees, a relatively lean headcount consistent with an infrastructure-focused engineering organization rather than a large services business. Fireworks AI is not a robotics hardware manufacturer; it is a horizontal AI model-serving and fine-tuning platform. Its relevance to the robotics and embodied-AI space comes from providing low-latency inference and custom-model training infrastructure that downstream developers -- including those building agentic software, multimodal systems and robotics-adjacent AI applications -- can use to deploy foundation models in production without operating their own GPU clusters.
Primary type & automation activities this supplier delivers:
Serverless and dedicated LLM inference, fine-tuning and multi-model routing (Fireworks Nexus) platform hosting open-source models such as DeepSeek, Kimi and Qwen.
Contact Fireworks AI
WEBSITE
https://fireworks.aiHEADQUARTERS
United States
Company Facts
Founded
2022
Primary Role
Software/Algorithm
Company Size
employees 50-100
Primary Region
North America
Annual Sales
-
Funding Stage
-
Funding Total
-
Moonshot AI has made its Kimi K3 model available through Fireworks AI on Microsoft's Foundry platform, enabling Azure customers to utilize the model via OpenAI-compatible APIs. This integration allows companies to transition open models from evaluation to production without the need for managing their own GPU infrastructure. The significance of this development lies in the streamlined access it provides to enterprise users, facilitating the deployment of AI models in a managed environment. With Kimi K3 listed among over 20 open models on Foundry, organizations can leverage the capabilities of Fireworks AI's inference engine alongside Microsoft's governance and billing features. Looking ahead, the focus will be on how enterprises adopt the Kimi K3 model within their operations. As Moonshot AI continues to expand its distribution channels, the emphasis will be on enhancing the accessibility and usability of AI solutions for businesses. No further timeline was disclosed at the time of publication.
TechNode.com Jul 30, 2026 News FeedThe NVIDIA Blackwell platform has gained significant traction among top inference providers, including Baseten, DeepInfra, Fireworks AI, and Together AI, achieving reductions in cost per token by as much as 10 times. Building on this success, NVIDIA has introduced the Blackwell Ultra platform, which aims to enhance these cost efficiencies further. This development reflects NVIDIA's commitment to advancing AI technology and providing more affordable solutions for inference tasks, thereby supporting the growing demand for cost-effective AI applications in various industries.
NvidiaNews Feb 16, 2026