Kimi K3 on Fireworks: Frontier Intelligence You Can Own

Fireworks Nexus

Don't ration intelligence

Drop frontier open models in the harnesses your engineers use, cutting your spend in half without sacrificing speed or quality.

Fireworks for Work
Cost Efficiency
33%

Savings per merged PR

Scale
40T

Tokens served daily

Speed
100+

Tokens/second available

Nexus is the fastest path to frontier open intelligence for code generation

Token bills are exploding at 20x YoY due to limited closed model competition. It's an impossible tradeoff: spend and ship, or ration tokens and stall. Nexus offers a better way.

FireConnect download command

Zero-friction setup

Developers set up in minutes with our CLI in the harnesses they use

  • One-command setup with FireConnect CLI.
  • Works with Claude Code, Codex, OpenCode, and more.
  • Authenticate with your preferred SSO provider for Fireworks keys
  • Choose open models or FireRouter— no proxy, no config surgery, fully reversible
Turn on FireRouter command

Intelligent routing

Developers set FireRouter as the model for optimal price-performance routing.

  • Each request is scored and routed to the optimal model for the task
  • Difficult tasks pass through to your closed provider on your key
  • Savings from open model routing and a 95%+ cache hit rate at half price per token
  • Tune routing preferences from max-intelligence to max-savings
Manage user limit command

Maximum control

Set user budgets and measure ROI on your workloads.

  • Default limit for the account, override it per user
  • Spend tracking by day, user, model, and API key
  • Fast and Priority options available for 100+ TPS and reliability
  • Measure blended cost per token, cost per merged PR, and more on your workloads

Frontier open models can now tackle 80%+ of engineering tasks

Faros AI + Fireworks
Production repositories

GLM-5.2 beat Opus 4.8 on quality, speed, and cost

211 engineering tasks · 12 repositories · 7 model-and-harness routes

0.568 judge score vs. 0.521

2.4X faster 321s vs. 775s

48% lower cost $0.92 vs. $1.76

What it proves:
Open models can outperform a closed frontier default on production-like engineering work - not just public benchmarks
Fireworks Evaluation
Agentic benchmark suite

K3 tied Fable on SWE, won more terminal tasks

≈1,030 tasks · 5 work families · Same harness used across tasks tested

92.4% SWE solve rate vs. 92.6%

11-7 solo wins across 89 terminal tasks

Up to 50X more cost-effective on long agentic loops

What it proves:
K3 delivers near-identical software-engineering quality, with strengths in security, cryptography, and long-horizon terminal work

Visibility, control, and better practices

Set your budgets. Track ROI. And give developers governed access to experiment with frontier open source models.

Set the ceiling

One default covers everyone and overrides handle the exceptions. A user who hits their limit can't make further requests until the billing period resets, unless you grant them an exception.

Measure ROI

Every dollar is queryable by day, model, user, and API key, with a raw per-event CSV export when you need the underlying rows. Measure savings, cost per merged PR, blended token rate, and more.

Keep the roster in sync

Developers authenticate through your identity provider with SSO enforcement by domain. JIT provisioning creates accounts on first sign-in. SCIM directory sync with Okta, Microsoft Entra ID, or Google Workspace adds and removes users as your directory changes.

Cultivate model citizens

Developers get fast, governed access to frontier open models to experiment and optimize. Check models and pricing with FireConnect model list or the Fireworks Dashboard to see exactly where you stand on spend before hitting a wall.

Set user budgets it the Fireworks dashboard

Take back control

Admins see every user's spend, limit, and override. Engineers see their own before they hit a wall. Set the ceiling once in Settings, firectl, or REST API.

Set up your harness in FireConnect

Bye bye, tokenmaxxing

Engineers maintain their tools and flow state with the FireConnect CLI. Setup in minutes. Create more token headroom with FireRouter, your model autopilot, or pin your own model choices.

FireConnect CLI Routing Preferences

Optimization on autopilot

FireRouter scores every request and takes the most cost efficient path between open and closed models with tunable preferences. Developers can pass --routing-preference when they enable a harness, or send the x-routing-preference header on individual calls.

Macroscope logo - dark

“Frontier models took us to product-market fit; now we're driving signal-to-noise further — more real bugs, fewer false positives. At our scale that's a training problem, and Fireworks makes it easy.”

Rob Bishop
Rob Bishop | Co-Founder at Macroscope
Sourcegraph

"Fireworks has been a fantastic partner in building AI dev tools at Sourcegraph. Their fast, reliable model inference lets us focus on fine-tuning, AI-powered code search, and deep code context, making Cody the best AI coding assistant. They are responsive and ship at an amazing pace."

Beyang Liu Testimonial
Beyang Liu | CTO at Sourcegraph
Gumloop

"Tokenmaxxing had a good run. But the teams that win with AI from here will be the ones getting more done for less, not the ones spending the most…After optimizing our harness for open weight models, we secretly swapped one of our most used internal agents from Opus 4.8 to GLM-5.2, and no one at the company noticed. We are now seeing cost savings of up to 72%."


Max Brodeur-Urbas | CEO at Gumloop
Macroscope logo - dark

“Frontier models took us to product-market fit; now we're driving signal-to-noise further — more real bugs, fewer false positives. At our scale that's a training problem, and Fireworks makes it easy.”

Rob Bishop
Rob Bishop | Co-Founder at Macroscope
Sourcegraph

"Fireworks has been a fantastic partner in building AI dev tools at Sourcegraph. Their fast, reliable model inference lets us focus on fine-tuning, AI-powered code search, and deep code context, making Cody the best AI coding assistant. They are responsive and ship at an amazing pace."

Beyang Liu Testimonial
Beyang Liu | CTO at Sourcegraph
Gumloop

"Tokenmaxxing had a good run. But the teams that win with AI from here will be the ones getting more done for less, not the ones spending the most…After optimizing our harness for open weight models, we secretly swapped one of our most used internal agents from Opus 4.8 to GLM-5.2, and no one at the company noticed. We are now seeing cost savings of up to 72%."


Max Brodeur-Urbas | CEO at Gumloop
Macroscope logo - dark

“Frontier models took us to product-market fit; now we're driving signal-to-noise further — more real bugs, fewer false positives. At our scale that's a training problem, and Fireworks makes it easy.”

Rob Bishop
Rob Bishop | Co-Founder at Macroscope

Leave failed experiments in the past

Tried self-hosting open models in the past? Open weights are clearing performance bars, but that wasn’t the only barrier to your success. Here's why this time is different:

Industry-leading open inference engine

Fireworks has built a fully disaggregated inference engine optimized at every layer, from custom kernels to memory management to adaptive caching. The platform now serves 40T+ tokens daily on a global scale with zero data retention, US-hosted only options, and certifications for SOC 2, ISO 27001, ISO 42001, HIPAA. We maintain commercial agreements for leading open frontier models and offer zero-day access with optimized performance.

Intelligent routing that's native to the inference stack

While it’s true you can configure a router in a weekend, a badly built ladder performs worse than no routing at all. Research by Arize proved naive escalation across ten models performed worse than every single model tested on its own. The judgment is the stack. For example, Fireworks holds a 95%+ cache hit rate on routine coding traffic, and cached input runs at half price. FireRouter factors hit rate into every decision.

Flexibility with your existing stack

Have a gateway? Keep it. Nexus has a documented LiteLLM Proxy integration. Use it for policy, fan-out, and fallbacks, and let FireRouter own the cost-versus-quality call on tasks inside it.

Interested in learning more?

Schedule a call with a forward-deployed engineer to discuss your code generation workloads and see Nexus in action.