Join us for our inaugural conference, Forge 2026

White Paper
Andrzej Test
Fireworks white paper featured image

How Cursor built Fast Apply using the Speculative Decoding API.

TL;DR

Millions of developers write code everyday to improve software systems at large bringing productivity to many business workflows, however, very few tools help them improve their own productivity.

With the advent of Generative AI, we see an emerging category of developer tooling built around Large Language Models (LLMs) especially for code generation. Some popular tools in this regard are Github Copilot, Sourcegraph, Phind, Continue, Blackbox, Codeium, Cognition, Factory, Aider., the list goes on…

Among them is a standout product, Cursor. Cursor is an AI-native IDE that helps developers write better code faster. Their core features include:

  1. Cursor’s Copilot++ is a more powerful AI-assistant which predicts your next edit taking in account of your recent change.
  2. Cmd/Ctrl-k, an instructed edit model that you can use to make changes to any region of code with natural language.
  3. A chat that sees your whole codebase and can “instantly apply” the changes to your code.

Our mission at Fireworks is to assist developers building innovative and mission critical Generative AI experiences at enterprise scale. In this blog, we will go through how Fireworks inference stack enabled Cursor to achieve 1000 tokens per sec using our Speculative Decoding API with low latency.

What Cursor is building?

There are many feature highlights in Cursor, but here are some that are making developers love them more.