DeepSeek-V4-Pro-0813 available now on Fireworks

Fireworks Blog

Header Image Do Open Models Perform Silent Reasoning Before They Respond

Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B

Three projects, one platform: A developer's winning streak with Fireworks
Case Studies
10/14/2024

Three projects, one platform: A developer's winning streak with Fireworks

Partnering with Meta: Bringing Llama 3.2 to Fireworks for Fine-Tuning and Inference
Model Releases
9/25/2024

Partnering with Meta: Bringing Llama 3.2 to Fireworks for Fine-Tuning and Inference

How Enterprises are using Multimodal Models in production with Fireworks
Case Studies
9/25/2024

How Enterprises are using Multimodal Models in production with Fireworks

Multi-LoRA: Personalize AI at scale and deliver the best experience for each customer and use case, with 100x cost-efficiency
Developer Experience
9/18/2024

Multi-LoRA: Personalize AI at scale and deliver the best experience for each customer and use case, with 100x cost-efficiency

FireOptimizer: Customizing latency and quality for your production inference workload
Model Releases
8/30/2024

FireOptimizer: Customizing latency and quality for your production inference workload

Build Your Own Flight Recommendation System using FastAPI, SerpAPI, and Firefunction
Developer Experience
8/29/2024

Build Your Own Flight Recommendation System using FastAPI, SerpAPI, and Firefunction

Building a RAG with Astro, FastAPI, SurrealDB and Llama 3.1
Developer Experience
8/14/2024

Building a RAG with Astro, FastAPI, SurrealDB and Llama 3.1

How Fireworks evaluates quantization precisely and interpretably
Developer Experience
8/1/2024

How Fireworks evaluates quantization precisely and interpretably

Introducing Llama 3.1 inference endpoints in partnership with Meta
Model Releases
7/23/2024

Introducing Llama 3.1 inference endpoints in partnership with Meta

Fireworks Raises $52M Series B to Lead Industry Shift to Compound AI Systems
Company News
7/11/2024

Fireworks Raises $52M Series B to Lead Industry Shift to Compound AI Systems

How Cursor built Fast Apply using the Speculative Decoding API
Developer Experience
6/23/2024

How Cursor built Fast Apply using the Speculative Decoding API

FireAttention V2: 12x faster to make Long Contexts practical for Online Inference
Model Releases
6/20/2024

FireAttention V2: 12x faster to make Long Contexts practical for Online Inference

Firefunction-v2: Function calling capability on par with GPT4o at 2.5x the speed and 10% of the cost=
Model Releases
6/17/2024

Firefunction-v2: Function calling capability on par with GPT4o at 2.5x the speed and 10% of the cost=

Announcing custom models and on-demand H100s with 50%+ lower costs and latency than  vLLM
Model Releases
6/3/2024

Announcing custom models and on-demand H100s with 50%+ lower costs and latency than vLLM

GPUs on-demand: Not serverless, not reserved, but some third thing
Developer Experience
6/3/2024

GPUs on-demand: Not serverless, not reserved, but some third thing

Code Generation with Large Language Models - Fireworks Take
Developer Experience
5/8/2024

Code Generation with Large Language Models - Fireworks Take

Doomed to Code: How we Teamed Up with Fireworks at MistralAI Hackathon to Conquer the Shores of Hell
Developer Experience
5/6/2024

Doomed to Code: How we Teamed Up with Fireworks at MistralAI Hackathon to Conquer the Shores of Hell

Partnering with Meta to bring Llama 3 to Firework’s inference and fine-tuning
Model Releases
4/18/2024

Partnering with Meta to bring Llama 3 to Firework’s inference and fine-tuning

Getting Started with Stability’s API Powered by Fireworks
Developer Experience
4/17/2024

Getting Started with Stability’s API Powered by Fireworks

Optimizing Retrieval Augmented Generation (RAG) with MongoDB Atlas and Fireworks
Developer Experience
3/21/2024

Optimizing Retrieval Augmented Generation (RAG) with MongoDB Atlas and Fireworks

multi-operation-fusions-MoE
Developer Experience
3/10/2024

Training-Inference Parity in MoE Models: Where Numerics Drift