DeepSeek-V4-Pro-0813 available now on Fireworks
Product
Solutions
Models
Pricing
Resources
Log In
Get Started
Fireworks Blog
Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B
Read More
Case Studies
Model Releases
Benchmarks
Partner Announcements
Developer Experience
Company News
Agentic
Use Cases
Multimodal
Training
Filters
Model Releases
3/8/2024
Fireworks launches fine-tuning service - Rapidly iterate on quality and scale to production through Fireworks inference
Model Releases
3/1/2024
Fireworks Platform Spring 2024 Updates
Model Releases
2/20/2024
FireFunction V1 - Fireworks’ GPT-4-level function calling model - 4x faster than GPT-4 and open weights
Model Releases
2/20/2024
Why do all LLMs need structured output modes?
Model Releases
1/18/2024
FireLLaVA: the first commercially permissive OSS LLaVA model
Developer Experience
1/8/2024
FireAttention — Serving Open Source Models 4x faster than vLLM by quantizing with ~no tradeoffs
Model Releases
12/20/2023
Fireworks Raises the Quality Bar with Function Calling Model and API Release
Model Releases
12/14/2023
Mixtral 8x7B on Fireworks: faster, cheaper, even before the official release
Developer Experience
11/3/2023
LLM Inference Performance Benchmarking (Part 1)
Model Releases
11/2/2023
New in Fireworks: Image-to-Image and ControlNet support for SSD-1B and SDXL!
Company News
10/27/2023
Fireworks.ai Achieves SOC 2 Type II and HIPAA Compliance
Model Releases
10/11/2023
Accelerating Code Completion with Fireworks Fast LLM Inference
Model Releases
10/2/2023
Fireworks.ai Now Available on LangChain Prompt Playground
Developer Experience
9/12/2023
Simplifying Code Infilling with Code Llama and Fireworks.ai
Developer Experience
8/29/2023
Speed, Python: Pick Two. How CUDA Graphs Enable Fast Python Code for Deep Learning
Model Releases
8/17/2023
Fireworks.ai: Fast, Affordable, Customizable Gen AI Platform
Developer Experience
7/12/2023
Multi-Query Attention is All You Need
Previous