Join us for our inaugural conference, Forge 2026

Fireworks Blog

Phylo and Fireworks

Phylo brings frontier AI to more scientists with open models on Fireworks

Phylo cut inference cost 60% while doubling users month-on-month, running Biomni Lab's long-horizon biology agents on open models on Fireworks.

Announcing custom models and on-demand H100s with 50%+ lower costs and latency than  vLLM
Model Releases
6/3/2024

Announcing custom models and on-demand H100s with 50%+ lower costs and latency than vLLM

GPUs on-demand: Not serverless, not reserved, but some third thing
Developer Experience
6/3/2024

GPUs on-demand: Not serverless, not reserved, but some third thing

Code Generation with Large Language Models - Fireworks Take
Developer Experience
5/8/2024

Code Generation with Large Language Models - Fireworks Take

Doomed to Code: How we Teamed Up with Fireworks at MistralAI Hackathon to Conquer the Shores of Hell
Developer Experience
5/6/2024

Doomed to Code: How we Teamed Up with Fireworks at MistralAI Hackathon to Conquer the Shores of Hell

Partnering with Meta to bring Llama 3 to Firework’s inference and fine-tuning
Model Releases
4/18/2024

Partnering with Meta to bring Llama 3 to Firework’s inference and fine-tuning

Getting Started with Stability’s API Powered by Fireworks
Developer Experience
4/17/2024

Getting Started with Stability’s API Powered by Fireworks

Optimizing Retrieval Augmented Generation (RAG) with MongoDB Atlas and Fireworks
Developer Experience
3/21/2024

Optimizing Retrieval Augmented Generation (RAG) with MongoDB Atlas and Fireworks

multi-operation-fusions-MoE
Developer Experience
3/10/2024

Training-Inference Parity in MoE Models: Where Numerics Drift

Fireworks launches fine-tuning service - Rapidly iterate on quality and scale to production through Fireworks inference
Model Releases
3/8/2024

Fireworks launches fine-tuning service - Rapidly iterate on quality and scale to production through Fireworks inference

Fireworks Platform Spring 2024 Updates
Model Releases
3/1/2024

Fireworks Platform Spring 2024 Updates

FireFunction V1 - Fireworks’ GPT-4-level function calling model - 4x faster than GPT-4 and open weights
Model Releases
2/20/2024

FireFunction V1 - Fireworks’ GPT-4-level function calling model - 4x faster than GPT-4 and open weights

Why do all LLMs need structured output modes?
Model Releases
2/20/2024

Why do all LLMs need structured output modes?

FireLLaVA: the first commercially permissive OSS LLaVA model
Model Releases
1/18/2024

FireLLaVA: the first commercially permissive OSS LLaVA model

FireAttention — Serving Open Source Models 4x faster than vLLM by quantizing with ~no tradeoffs
Developer Experience
1/8/2024

FireAttention — Serving Open Source Models 4x faster than vLLM by quantizing with ~no tradeoffs

Fireworks Raises the Quality Bar with Function Calling Model and API Release
Model Releases
12/20/2023

Fireworks Raises the Quality Bar with Function Calling Model and API Release

Mixtral 8x7B on Fireworks: faster, cheaper, even before the official release
Model Releases
12/14/2023

Mixtral 8x7B on Fireworks: faster, cheaper, even before the official release

LLM Inference Performance Benchmarking (Part 1)
Developer Experience
11/3/2023

LLM Inference Performance Benchmarking (Part 1)

New in Fireworks: Image-to-Image and ControlNet support for SSD-1B and SDXL!
Model Releases
11/2/2023

New in Fireworks: Image-to-Image and ControlNet support for SSD-1B and SDXL!

Fireworks.ai Achieves SOC 2 Type II and HIPAA Compliance
Company News
10/27/2023

Fireworks.ai Achieves SOC 2 Type II and HIPAA Compliance

Accelerating Code Completion with Fireworks Fast LLM Inference
Model Releases
10/11/2023

Accelerating Code Completion with Fireworks Fast LLM Inference

Fireworks.ai Now Available on LangChain Prompt Playground
Model Releases
10/2/2023

Fireworks.ai Now Available on LangChain Prompt Playground