Announcing our Series D and $1B ARR

Model Library
/Fireworks AI/Ling 3 Flash
Fireworks Logo Mark

Ling 3 Flash

Ready
model path:accounts/fireworks/models/ling-3-flash

Ling 3 Flash is currently planned as a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

Ling 3 Flash API Features

On-demand Deployment

Docs

On-demand deployments allow you to use Ling 3 Flash on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits.

Metadata

State
Ready
Created on
7/23/2026
Kind
Base model
Provider
Fireworks AI

Specification

Calibrated
No
Mixture-of-Experts
Yes
Parameters
124B

Supported Functionality

Fine-tuning
Not supported
Serverless
Not supported
Context Length
256k tokens
Function Calling
Supported
Embeddings
Not supported
Rerankers
Not supported
Support image input
Not supported