Muse Glimmer 30B is a dense causal language model distilled from Muse Spark and purpose-built for autonomous agentic work. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a ~1.8B ViT-G/14 perception encoder, supporting interleaved text and image input, a 131K+ context window, and selectable reasoning strength (low through xhigh). Trained on data from over 100 languages, Muse Glimmer performs strongly for its size class on agentic benchmarks including MCP Atlas, DeepSearch QA, Gaia2 and SWE-Bench Pro, and is released under Apache 2.0.
Muse Glimmer 30B can be customized with your data to improve responses. Fireworks uses LoRA to efficiently train and deploy your personalized model
On-demand deployments allow you to use Muse Glimmer 30B on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits.
Only on serverless, starting on September 25, 2026. On-demand deployment will continue to be available. Migrate serverless workloads to Nemotron Lightning 3.5 30B A3B.