DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
DeepSeek-V4-Flash-0731 can be customized with your data to improve responses. Fireworks uses LoRA to efficiently train and deploy your personalized model
On-demand deployments allow you to use DeepSeek-V4-Flash-0731 on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits.
Only on serverless, starting on September 25, 2026. On-demand deployment will continue to be available. Migrate serverless workloads to DeepSeek V4.1 Flash.
A Mixture-of-Experts model from DeepSeek, released 2026-07-31. It is the official release of DeepSeek-V4-Flash, superseding the preview with enhanced agentic capabilities and an attached speculative decoding module, sharing the structure of DeepSeek-V4-Flash-DSpark.
Coding agents, terminal automation, and multi-step agentic workflows at long context. DeepSeek reports 82.7 on Terminal Bench 2.1, 70.3 on Toolathlon-Verified, and 76.7 on Cybergym.
1,048,576 tokens, set as max_position_embeddings and reached through YaRN rope scaling at factor 16 over a 65,536-token base.
Fireworks serves a 1040k-token context window, available on both the serverless endpoint and on-demand deployments.
DeepSeek recommends a temperature of 1.0. For local deployment it recommends top_p of 0.95 for agentic scenarios and 1.0 otherwise.
Fireworks defaults max_tokens to 2,048 and allows generation up to the full context window. DeepSeek recommends capping output at 384K tokens at high or max reasoning effort.
Yes to both. Streaming is set with stream, with usage in the final chunk, and function calling uses OpenAI-compatible JSON Schema tool definitions.
304 billion total parameters as listed by Fireworks, with 13 billion active per token across 256 routed experts and 1 shared expert. DeepSeek publishes 284 billion for the base model.
Yes. Two dedicated shapes run at 262,144-token context: full-parameter on 8×B300 and LoRA on 4×B300, with managed SFT and DPO on the LoRA shape.
Fireworks reports prompt_tokens, completion_tokens, and total_tokens in the response usage object. Serverless billing prices input, cached input, and output separately at $0.14, $0.028, and $0.28 per 1M tokens.
Serverless ceilings for this model's size tier default to 64.8M total prompt TPM, 16.2M uncached prompt TPM, and 648k generated TPM, adaptive per account and model. On-demand deployments carry no rate limits.
Fireworks announces serverless model deprecations in advance, following its serverless model lifecycle policy. For long-term version stability it recommends on-demand deployments.
The MIT License, which permits commercial use, modification, and redistribution subject to retaining the copyright and license notice.
No. Fireworks operates zero data retention by default and does not log prompt or generation data for open models without opt-in. The Responses API is the exception, retaining data 30 days when store=True.