Join us for our inaugural conference, Forge 2026

Get listed on the Specialized Intelligence Index

Contribute a real-work benchmark from your domain. Fireworks reviews the evaluation, runs it across leading open and closed models, and publishes results with your approval.

What you get

  • Benchmark the frontier. See how leading open and closed models perform on your evaluation.
  • We cover the compute. Fireworks runs the model roster at no cost to you.
  • Keep your eval private. Tasks, prompts, graders, and trajectories remain private unless you choose to share them.
  • Publish credible comparisons. Results include quality, cost, and latency under controlled evaluation conditions.
  • Improve what falls short. Fireworks Lab can post-train models against identified failure modes and re-evaluate them on the same benchmark.

Loading...