
Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default.
→ Use Priority for max reliability during congestion; priced at +25% from standard rates.
→ Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast
→ Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us
Fine-tuningDocs | Kimi K3 can be customized with your data to improve responses. Fireworks uses LoRA to efficiently train and deploy your personalized model |
ServerlessDocs | Kimi K3 is available via Fireworks' serverless API, where you pay per token. There are several ways to call the Fireworks API, including Fireworks' Python client, the REST API, or OpenAI's Python client. |
On-demand DeploymentDocs | On-demand deployments allow you to use Kimi K3 on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits. |
Run queries immediately, pay only for usage