Announcing our Series D and $1B ARR

Fireworks Blog

Headline Image Showing Why Routing Kimi K3 and Fable Outperforms the Rest

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

We ran both Kimi K3 and Fable 5 through ~1,000 agentic benchmark tasks. They tie on the overall top-line numbers, but specialize beneath the surface. K3 outperforms terminal and dev tooling, while Fable leads on web and multi-language tasks. Most importantly, we demonstrate that efficient routing between them improves overall accuracy and dramatically reducing token spend.

Best practices for multi-turn RL
Developer Experience
12/10/2025

Best Practices for Multi-Turn RL

Turn Your LLM into a Calibrated Classifier for $2
Developer Experience
12/4/2025

Turn Your LLM into a Calibrated Classifier for $2

Unlock Advanced Reasoning with NVIDIA Nemotron Nano 2 Models on Fireworks AI
Model Releases
12/2/2025

Unlock Advanced Reasoning with NVIDIA Nemotron Nano 2 Models on Fireworks AI

Fireworks Expands AWS Alliance: Strategic Collaboration Agreement
Partner Announcements
11/24/2025

Fireworks Expands AWS Alliance: Strategic Collaboration Agreement + GenAI Competency

Eval Protocol: RL on your agents, in any environment
Developer Experience
11/20/2025

Eval Protocol: RL on your agents, in any environment

Fireworks ISO Certifications
Company News
11/19/2025

Fireworks Achieves Triple ISO Certification, giving Enterprises Full Control and Trust in AI at Scale

50 Trillion Tokens Per Day The State of Agent Environments
Developer Experience
11/19/2025

50 Trillion Tokens Per Day: The State of Agent Environments

Fireworks RFT: Build AI Agents with fine-tuned open models that outperform frontier closed models
Developer Experience
11/10/2025

Fireworks RFT: Build AI agents with fine-tuned open models that outperform frontier closed models

RADPAIR and Fireworks Unlock Smarter Radiology Workflows
Case Studies
11/9/2025

Modernizing Healthcare with AI: How RADPAIR and Fireworks Unlock Smarter Radiology Workflows

Vercel and Fireworks Partnership
Case Studies
11/3/2025

40X Faster, and Smarter Outputs: How Vercel Turbocharged their Code Fixing Model with Open Models, Speculative Decoding and Reinforcement Fine Tuning on Fireworks

Genspark’s Deep Research Agent Outperforms a Frontier Closed Model in Quality and Tool Calls using Fireworks Reinforcement Fine Tuning, Achieving a 50% Cost Reduction
Case Studies
10/31/2025

Genspark’s Deep Research Agent Outperforms a Frontier Closed Model in Quality and Tool Calls using Fireworks RFT, Achieving a 50% Cost Reduction

Series C
Company News
10/28/2025

We raised $250M To Help Enterprises Own Their AI

Deploy NVIDIA Nemotron Nano 2 VL on Fireworks
Model Releases
10/27/2025

Accelerate your Vision Pipelines with the new NVIDIA Nemotron Nano 2 VL Model on Fireworks AI

Deployment Shapes One Click Deployment Configured for You
Developer Experience
10/23/2025

Deployment Shapes: One-Click Deployment Configured For You

fireworks amd
Partner Announcements
10/20/2025

Fireworks and AMD partner to power the next gen of AI infrastructure on AMD Instinct™ GPUs

LLM on the edge: Model picking with Fireworks Eval Protocol + Ollama
Developer Experience
10/15/2025

LLM on the edge: Model picking with Fireworks Eval Protocol + Ollama

Announcing Embeddings  and Reranking  on Fireworks AI
Model Releases
10/9/2025

Announcing Embeddings and Reranking On Fireworks AI

Deep-Dive into LLM Fine Tuning
Developer Experience
10/6/2025

Deep-Dive into LLM Fine-Tuning

Production-Ready AI Agents with Optimized Inference with AWS AgentCore
Developer Experience
10/2/2025

Production-Ready AI Agents with Optimized Inference with AWS AgentCore

Fireworks for Startups
Company News
10/1/2025

Launching Fireworks for Startups Program!

image
Developer Experience
9/22/2025

Traces Are All You Need (to rank LLMs)