Contribute a real-work benchmark from your domain. Fireworks reviews the evaluation, runs it across leading open and closed models, and publishes results with your approval.
What you get
•Benchmark the frontier. See how leading open and closed models perform on your evaluation.
•We cover the compute. Fireworks runs the model roster at no cost to you.
•Keep your eval private. Tasks, prompts, graders, and trajectories remain private unless you choose to share them.
•Publish credible comparisons. Results include quality, cost, and latency under controlled evaluation conditions.
•Improve what falls short. Fireworks Lab can post-train models against identified failure modes and re-evaluate them on the same benchmark.