
If you run engineering at any real scale, you've already had this conversation: someone in finance asks what the AI cost is going to be next quarter, and the honest answer is that nobody knows.
When Forbes reported that Uber had spent its entire 2026 AI budget by April, the news made global headlines. They rolled out Claude Code to roughly 5,000 engineers in December and watched agentic adoption climb from about a third of engineers to more than four-fifths inside two months. The CTO reported burning $1,200 in a single two-hour session. Uber’s COO said publicly that the link between that spend and shipped customer value is hard to draw.
But something’s been quietly changing: open models are now reliable for the vast majority of real tasks. We found most organizations were running routine work at frontier prices, but due to the operational complexity of running open-weight models at scale, it’s been hard for them to switch. So we took all the background work out of running the best, curated open models like Kimi-K3 and GLM-5.2 at scale, and made it something you drop straight into your existing workflow. We’ve been already testing Fireworks Nexus with teams like Notion and Doximity where the preliminary results show that Fireworks Nexus cuts a third off per merged PR costs, and has a blended token rate roughly a quarter of the closed model labs.
Fireworks Nexus connects the AI tools your developers already use to a managed layer of open-weight models, with enterprise controls and intelligent routing. Engineering teams gain visibility and control over the AI powering their workflows while lowering costs and maintaining the developer experience they already know.
It is composed of three main components:
Fireworks Nexus gives engineering and IT teams centralized control over how AI is used across the organization. Set budgets at the team or company level, track ROI across models and tools, and enforce policies from a single place. Behind the scenes, every request runs on the same production inference platform trusted by leading AI-native companies, with US-hosted endpoints, zero data retention, and enterprise-grade performance across 20 global data centers.

FireConnect is a one-line install that maps appropriate models based on harness configurations intuitively. Keep using your current tools and harnesses like Claude Code, Codex, OpenCode. It ensures developers face no disruption, while benefiting from optimized infrastructure including higher cache rates that can drive material cost savings. FireConnect works through our Fireworks Serverless APIs that are Anthropic and OpenAI API compatible, so most tools connect with a base URL and a model ID. To get started, FireConnect is open sourced under Apache 2.0 and can be installed from the Fireworks Dashboard in a single command.

We built a custom routing endpoint for organizations with existing AI contracts. A custom trained model scores each request's difficulty. Requests that are routine tasks go to a cost-effective open-weight model served by Fireworks, difficult tasks pass through to your existing provider on your own key, which is never stored server-side. It serves as an intermediary layer that dynamically routes tasks based on complexity and optimal caching. This typically delivers a 3–5x cost reduction without sacrificing quality. This feature is in research preview and today routes between Claude Opus 5 and GLM 5.2, so the pass-through path needs an Anthropic key. Want to stay entirely on open weight models? We also route between K3 and GLM 5.2


The value of Fireworks Nexus isn’t just lower token prices. It’s getting the same or better engineering outcomes while spending less.
Independent evaluations back this up. Faros ran 211 real engineering tasks from customer repositories across seven model-and-harness combinations and published the full methodology. Claude Code running on GLM-5.2 slightly outperformed Claude Code on Opus 4.8 while costing roughly half as much per completed task. Arize evaluated 2,400 agent runs and found that frontier-priced models offered little advantage on routine work, while intelligent routing preserved their strengths where they mattered most. Their evaluation harness is open source, so you can run the same analysis on your own workloads.
Fireworks Nexus gives engineering organizations control over the intelligence powering the AI tools they already use. Instead of being locked into a single provider’s pricing, models, and roadmap, you can choose the best model for every task, manage spend centrally, and continuously measure quality as new models emerge.
Ready to see what Nexus can do for your engineering team? Request a demo and we’ll show you how it works on your own workloads.