Qwen3.8-Flash-Next is an experimental Qwen4-architecture preview: a multimodal MoE with 125B parameters (6B active) plus 51B n-gram embeddings and native 262K context.
On-demand DeploymentDocs | On-demand deployments allow you to use Qwen3.8 Flash Next FP4 on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits. |