Measure models on real work
See how open, closed, and specialized models perform on domain-specific tasks drawn from production workloads.
Benchmarks built by practitioners
Evaluate models against the tasks, constraints, and quality standards that define acceptable work in each domain.
Choose models with confidence
Compare quality, cost, and task duration on the work that matters to your organization, not generalized intelligence alone.
{{ ind.chartAlertTag }}{{ ind.chartAlert }}
{{ ind.chart.reasoningNote }}
{{ ind.chart.singleMetricNote }}
{{ ind.chartNote }}
Specialized model highlights
{{ highlightAtLabel }}{{ highlightTotal }}
{{ hl.claim }}
{{ hl.body }}
{{ hl.note }}
{{ bar.valueStat }} {{ bar.valueLabel }}
Submit your model and/or benchmark for consideration. Don’t have one yet? Fireworks can help.
This methodology is versioned, published alongside the index, and inherited by every benchmark run unless a benchmark or model row records an exception.
The Specialized Intelligence Index measures how models perform on domain benchmarks under one standardized serving configuration and one harness per benchmark. It does not measure the maximum achievable performance on a benchmark, which a team optimizing one model with a purpose-built harness will typically exceed.
The index does not measure general capability and is not a substitute for a generalist index. Benchmarks are contributed by domain partners that use them in production or drawn from recognized public sources. Fireworks runs open, closed, and custom models against benchmarks built by their owners; it does not design the benchmarks.
Open-weight models are selected from top performers on our internal and external evaluations. Closed models run alongside them through first-party APIs, and partner fine-tunes run on their own domain benchmarks.
Scores on benchmarks listed as “Reproducible by Provider” are provided to us directly by the provider. On all other benchmarks, we generally collect three runs and report the average (avg@3), along with a 95% confidence interval; benchmark-native aggregations such as pass@3 are named on the card. Where a benchmark’s native metric is not pass/fail, its card names the metric. Metrics are never averaged across benchmarks.
Score carries its 95% margin beside it as ±x.x, with the full interval in the cell's tooltip. Cost and duration name the denominator they were measured against: Cost / task and Duration / task by default, and / trial where the source measured one trial of a repeated task. Both state their exclusions in the table footer. Duration covers the full end-to-end execution time, excluding judge time.
Default profile SII-2026.09:
Fireworks-executed evaluations use open-source harnesses pinned to versioned snapshots. The same harness, orchestrator, tools, and tool definitions apply to every model on a benchmark.
Reproducibility identifies who can rerun a score, not whether a rerun will produce an identical result.
The selected “Rank by” metric first orders eligible models within each benchmark. Higher score ranks first; lower cost or duration ranks first. Exact benchmark-score ties receive the average of their occupied positions.
Benchmark visibility is the owner’s decision. Some partners publish their benchmarks; others keep them private. For private benchmarks, scores are published, while evaluation scripts, prompts, grading logic, and trajectories are not.
{{ part.body }}
{{ bench.runCountNote }}
Hard RCA ranks · 324 verdicts per reportable run · 2–3 runs per model
{{ dfbenchNote }}