Kimi K3 on Fireworks: Frontier Intelligence You Can Own

Blog
Kimik3 On Fireworks

Kimi K3 on Fireworks: Frontier Intelligence You Can Own

Benchmarks Showing Opus 5 vs Kimi K3

US-Hosted, Zero Data Retention, and Day-0 Training

Kimi K3 is open-weight today, and available Day-0 for both inference and training on Fireworks. It delivers frontier-level intelligence, the top open model in the world, at a fraction of closed-model cost. It is #1 at frontend code, strong at writing, and it reliably finishes the long, multi-step agentic tasks where other models stall.

Michele Catasta, the President of Replit, told us,

"Kimi K3 on Replit Design exceeded our expectations. Unprecedented product design and UI capabilities, at a fraction of the cost we've come to expect from frontier models."

This is a turning point. Open models have crossed the line where they match closed frontier quality on real work while costing far less. That changes the default: start with an open model like Kimi K3 for the bulk of your tasks, see exactly what you are spending, and route to specialized intelligence only where a task demands it.

With Kimi K3 on Fireworks, you own your roadmap. Training a 3T-class model can start from your laptop, no infrastructure setup, no deployments to bring up. Just start a session and send tokens. And we keep your private data private. With Fireworks, specialized intelligence becomes a moat that compounds with every cycle.

That ownership comes with the guarantees regulated teams need: US-hosted inference and zero data retention, on an API that stays out of your way.

Michael Haines, Product Lead at Mercor, put it this way:

"Evaluating new models against our benchmarks usually requires trade-offs between performance, cost and scale, but Kimi K3 on Fireworks delivered across the board. The performance is top-tier, the serverless scaling is seamless, and the Zero Data Retention assurance gives us full confidence at scale."

That ownership changes the economics. What matters is your cost per finished task, set by how often the model succeeds on the first try and how many steps it takes. Tailoring Kimi K3 raises the success rate and shortens the trajectory, so budget shifts from generic tokens to the differentiation that makes your company unique.

Kimi K3 Rivals Anthropic and OpenAI’s Top Models

Kimi K3 is the first open model to reach 2.8 trillion parameters, providing frontier-level reasoning that rivals closed models like Fable 5, Opus 5 and GPT 5.5. Kimi K3 is designed to handle the most demanding use cases in software development, cybersecurity, knowledge work, multi-modal and visual work. What makes K3 credible is that most of the standout results come from independent benchmarks:

  • #1 World Rank in Front-End Code: K3 topped Arena's Frontend Code Arena at 1,679 points, surpassing Fable 5. This makes it the first open model to sit ahead of every closed one. It also claims top spot on Vercel's own Next.js evaluation suite.
  • Frontier Full Stack Coding: It ranked #3 on DeepSWE (Datacurve), only behind Fable 5 and GPT-5.6 Sol.
  • Top Open Model for General Intelligence: It is ranked #3 in the world on the Artificial Analysis Intelligence Index at 57, comparable to Opus and GPT-5.5-class models.
  • #1 for writing in editorial voice on 2840 Elo: surpassing Claude Fable 5. That is a jump from #21 to #1 over its predecessor, Kimi K2.6.
  • Sets a New Standard for Legal Analysis: K3 achieved a 26.7% all-pass rate on Harvey’s LAB legal benchmark, outperforming Fable 5 by nearly 2x.
  • Unmatched Performance for Customer service: On Sierra’s Tau3-Banking benchmark, K3 scored 33.4%, narrowly beating GPT 5.6 Sol.
  • Top Tier on Multimodal Intelligence: It ranks in the top two models globally for complex chart, screenshot, and dense document analysis (CharXiv and Zerobench).
  • High-End Cybersecurity: Similar recall as GPT 5.5 on Vercel’s private cybersecurity eval at a much lower price.
  • Top Tier for Autonomous Game Development: The developer community is buzzing about the new frontier for Agentic Game Development, from Night Rider to Fluppy Bird
Benchmark Chart
Figure 1: Third-Party Frontier Performance Benchmarks on Kimi K3

Last week, our research team shared how K3 matches Fable on quality for a fraction of the cost in long-running agentic workflows. By routing tasks between K3 and Fable, you can unlock the best possible performance for your specific needs.

Opus 5 vs Kimi K3 at a Glance

Following last week’s Opus 5 release, our team benchmarked it head-to-head against Kimi K3. We found Kimi K3 delivers matching performance at up to 5x better cost efficiency per task. Independent benchmarks like Vals Index found the exact same thing in the quality of the K3 against Opus.

TaskModelAccuracy$/taskTurns/task
SWE (480) Kimi K392.7%$0.5255.6
Opus 594.8%$1.0537.9
Algorithmic(100)Kimi K388.0%$0.0643.4
Opus 588.0%$0.1764.1
Terminal (83)Kimi K381.9%$0.3511.6
Opus 585.5%$1.6118.7

Kimi Delta Attention: Faster Tokens at a Lower Cost for Large Context Workloads

Hosting K3 yourself isn’t practical for most engineering teams. Moonshot recommends running it on a cluster of 64+ accelerator supernodes(GPUs). This means huge upfront hardware spend, and endless money wasted on idle GPUs. Fireworks lets you skip the cluster headaches, and instead of burning budget on dedicated GPU reservations, you pay per token and get frontier-level reasoning wherever you need it.

K3 is not just bigger; it is more efficient by design. Moonshot improved the design by changing how information flows across the model. Our team at Fireworks put together custom kernels for Kimi Delta Attention, FP4 MoE kernels, and adapted decode kernels to unlock K3’s full performance. The model architecture brings a new Kimi Delta Attention architecture and refined attention residuals. They scaled up the Mixture of Experts (MoE), activating 16 out of 896 experts with the Stable Latent MoE. This approach delivered 2.5 times the scaling efficiency compared to Kimi K2. This result is a model that delivers frontier reasoning for large, long-horizon tasks, while remaining cost-effective for high-volume production workloads.

The Specs

SpecKimi K3Kimi K2.7GLM 5.2
TierFrontierCodingCoding
Modalities Text and Vision Text and Vision (MoonViT) Text
Total Parameters2.8T1T MoE753 MoE
Input Context Length (Tokens) 1M 256k1M
Active Parameters104B32B40B
Activation Rate3.7%3.20%5.31%
# of Experts896384256

Run Kimi K3 the way you want on Fireworks

You can use Kimi K3 on Fireworks interchainably with most of your AI workloads. Our serverless platform simplifies using the model to just drop in an API key. You don’t have to think about GPU management or API compatibility.

Fireworks has increased default rate limits by over 10x for Serverless inference, launched Priority Serverless for better reliability, and optimized Fast Serverless for better token speed. The best part is you don’t need to rent GPUs for all of these options: just pay for the tokens you use like you do with Claude and Open AI.

“When we did a bakeoff with other providers, Fireworks won simply because it worked consistently. Whenever we deploy any model, it works the first time. No tuning, no fiddling. What I don't want is getting stuck in a 3-week development cycle trying to make a model work.”
— Travis Rehl, CTO Innovation Solutions

For work you can delay, Batch mode runs jobs at off-peak times for 50% off. It's how teams use K3's vision reasoning to process whole folders of images, captioning, labeling, or cropping, and the same savings apply to anything batchable, like running evals.

Built for Regulated Industries with US-only Serverless and Zero Data Retention

Today, we’re also launching US-only Serverless endpoints, starting with Kimi K3. Over the coming days and weeks, we are adding other US-only endpoints for the most popular models on Fireworks. Financial services, healthcare, and other regulated companies that require US data residency can now sign up and hit an endpoint with Zero Data Retention and US located inference.

For these security-sensitive workloads where data privacy is non-negotiable, Fireworks operates on Zero Data Retention by default. We never log or store your prompt or generation data for open models without your explicit opt-in, giving your team complete data control and peace of mind.

But why stop there? Making K3 yours starts from your laptop

Fine-tuning a frontier model is usually a heavy lift. Now in private preview: Fireworks Serverless Training.

It is a one-line code change from our reserved training SDK. Write a standard Python loop, point it at our Serverless Training API, and start iterating in seconds. With the K3 launch we are opening private preview access to a shared GPU pool, so there is no GPU reserved capacity to provision and you pay per token instead of by the reserved hour.

Want the full walkthrough? Part two takes you through everything you need to know to train K3: what LoRA is, how to shape the reward, when a small adapter is enough, and what it costs. Click Here

K3 is on Fireworks now. Serve it, fine-tune it, and put it into production.