DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.
ServerlessDocs | DeepSeek V4.1 Flash is available via Fireworks' serverless API, where you pay per token. There are several ways to call the Fireworks API, including Fireworks' Python client, the REST API, or OpenAI's Python client. |
On-demand DeploymentDocs | On-demand deployments allow you to use DeepSeek V4.1 Flash on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits. |
Run queries immediately, pay only for usage