# Fireworks Specialized Intelligence Index

> Native benchmark scores, cost, and duration for open, closed, and specialized language models across domain benchmarks contributed by industry practitioners.

Machine-readable: [JSON](https://fireworks.ai/specialized-intelligence-index/api/index.json) · [llms.txt](https://fireworks.ai/specialized-intelligence-index/llms.txt)

Sources checked Sep 21, 2026. Generated 2026-09-22.

Units: each board keeps its native headline metric; percentage scores use a 0–100 scale, Diagnosis Elo is a relative rating, cost is US dollars per task except τ³-Banking (per trial) and Genspark Slides (per deck), and duration is wall-clock seconds per task except τ³-Banking and FrontierSWE V2, which report seconds per trial.

## Summary

- On dfbench, contributed by depthfirst, Mythos 5† (Anthropic) ranks first with 69% score (detection recall).
- On PWNBench v0.1, contributed by Novee, Grok 4.6 (xAI) ranks first with 63.31% score (F0.5), $330.65 per task.
- On CyberGym, DeepSeek V4.1 Flash (DeepSeek) ranks first with 94.82% score (pass@3), $6.53 per task, 1353 seconds per task.
- On LAB, contributed by Harvey, Grok 4.6 (xAI) ranks first with 22.08% score (all-pass), $4 per task, 2160 seconds per task.
- On LAB: Contracts, contributed by Harvey, Grok 4.6 (xAI) ranks first with 10% score (all-pass), $2.19 per task, 2160 seconds per task.
- On APEX Agents: Corporate Law, contributed by Mercor, GPT-6 Astra (OpenAI) ranks first with 87.67% score (avg@4), $4.42 per task.
- On Redline Bench, contributed by Crosby, GPT-6 Astra (max) (OpenAI) ranks first with 62.84% score (avg@3), $9.26 per task, 2286.4 seconds per task.
- On Big Finance Benchmark, contributed by Rogo, GPT-6 Astra (OpenAI) ranks first with 54.6% score (avg@3), $0.303 per task.
- On DuetBench–Diagnosis, contributed by Decagon, Claude Opus 5 (Anthropic) ranks first with 1034.9 score (Diagnosis Elo), $3.86 per task, 326.3 seconds per task.
- On τ³-Banking, contributed by Sierra, GLM-5.3 (Z.ai) ranks first with 51.63% score (mean trial reward), $0.4852 per trial, 281.6 seconds per trial.
- On τ³-Voice, contributed by Sierra, gpt-live-1 (OpenAI · backend: gpt-6 astra (medium) · v1.0) ranks first with 81.72% score (Pass@1), 2.5 seconds per task.
- On FrontierSWE V2, contributed by Proximal, GPT-6 Astra (OpenAI) ranks first with 65.5% score (mean@5), $5148.25 per task, 43560 seconds per trial.
- On APEX-SWE, contributed by Mercor, Claude Opus 5 (Anthropic) ranks first with 63.75% score (avg@4), $21.93 per task.
- On DeepSWE v1.1, Claude Opus 5 (Anthropic) ranks first with 72.27% score (avg@3), $6.18 per task, 1309 seconds per task.
- On MacroscopeBench, contributed by Macroscope, GPT-6 Astra · max (OpenAI) ranks first with 77.96% score (avg@3), $7.87 per task, 228.9 seconds per task.
- On ORCA Benchmark, contributed by Traversal, GPT-6 Astra (OpenAI) ranks first with 50.21% score (Hard RCA, avg@3), $4 per task, 1143.7 seconds per task.
- On Bedside Bench, contributed by Doximity, Doximity Ask V7 (Doximity) ranks first with 95.4% score (avg@3).
- On APEX-1: General Practitioner, contributed by Mercor, Claude Opus 5 (Anthropic) ranks first with 72.12% score (avg@4), $1.97 per task.
- On HealthBench Professional, GPT-6 Astra (OpenAI) ranks first with 64.43% score (avg@8), $0.0533 per task, 26.5 seconds per task.
- On Genspark Slides Benchmark, contributed by Genspark, Claude Fable 5.1 (Anthropic) ranks first with 83.59% score.

## Benchmarks

### dfbench (depthfirst)

Domain: Security. Status: published.

A held-out evaluation of three jobs a security engineer actually does: finding vulnerabilities in a codebase, judging whether a reported finding is real, and working out what a diff changed. Detection and differential analysis run on separate task sets with their own ground truth, so the two are scored and costed independently rather than averaged into one number.

Source: https://depthfirst.com

| Rank | Model | Provider | Score (detection recall) (percent) | Differential-analysis recall (percent) | Precision (percent) | F1 (percent) | Cost / task (USD) | Detection cost / task (USD) | Differential-analysis cost / task (USD) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Mythos 5† | Anthropic | 69 | — | 24.5 | 36.2 | — | 99.19 | — |
| 2 | GPT 5.6 Sol | OpenAI · xhigh | 65.7 | 75.3 | 18 | 28.3 | 27.02 | 43.37 | 10.67 |
| 3 | dfs-large1 | depthfirst | 62.2 | 75.6 | 18.3 | 28.3 | 4.055 | 6.77 | 1.34 |
| 4 | DeepSeek v4.1 Flash | DeepSeek · high | 57.3 | 70.1 | 36.4 | 44.5 | 1.2 | 1.69 | 0.71 |
| 5 | Grok 4.5 | xAI · high | 57.2 | 65.1 | 29.5 | 38.9 | 5.54 | 7.7 | 3.38 |
| 6 | GPT 5.6 Luna | OpenAI · xhigh | 52.4 | 71 | 23.2 | 32.2 | 1.575 | 2.53 | 0.62 |
| 7 | Kimi K3 | Moonshot AI · max | 48 | 60.3 | 30.1 | 37 | 7.72 | 9.1 | 6.34 |
| 8 | Opus 5 | Anthropic · med | 47.8 | 71.3 | 40.5 | 43.8 | 13.685 | 21.89 | 5.48 |
| 9 | Muse Spark 1.3 | Meta · xhigh | 47.4 | 72.8 | 34.1 | 39.7 | 6.68 | 10.24 | 3.12 |
| 10 | Gemini 3.8 Flash | Google · high | 44.2 | 63.8 | 41.4 | 42.8 | 4.8 | 8.03 | 1.57 |
| 11 | Qwen 3.8 | Alibaba · max | 40.8 | 72.7 | 29.6 | 34.3 | 5.45 | 7.91 | 2.99 |
| 12 | GLM 5.2 | Z.ai · xhigh | 40.7 | 61.4 | 28.6 | 33.6 | 5.88 | 9.72 | 2.04 |

Note: † Run with Claude Security.

### PWNBench v0.1 (Novee)

Domain: Security. Status: published.

Greybox pentesting against realistic web applications: the agent probes a running target and files the vulnerabilities it finds. Reports are scored on precision and recall against a known ground truth, weighted by severity, alongside the API cost per task of the attempt. Every model is run at several attempt budgets so coverage and spend can be read together.

Source: https://novee.security

| Rank | Model | Provider | Score (F0.5) (percent) | Precision (percent) | Recall (percent) | F0.5 (medium and above) (percent) | Precision (medium and above) (percent) | Recall (medium and above) (percent) | Cost / task (USD) | Duration / task (seconds) | Standard error (percentage points) | Attempt budget (k) (count) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Grok 4.6 | xAI | 63.3132 | 83.26 | 38.37 | 67.0087 | 80.14 | 48.59 | 330.65 | — | — | 3 |
| 2 | Opus 5 | Anthropic | 59.9613 | 65.58 | 50.56 | 61.5009 | 68.85 | 49.24 | 1399.69 | — | — | 3 |
| 3 | Opus 4.8 | Anthropic | 55.5876 | 74.55 | 28.99 | 61.8151 | 83.06 | 36.97 | 564.34 | — | — | 3 |
| 4 | GPT-5.6 Sol | OpenAI | 54.0096 | 74.93 | 31.21 | 52.8045 | 69.25 | 32.25 | 571.05 | — | — | 3 |
| 5 | DeepSeek V4.1 Flash | DeepSeek | 49.1864 | 52.14 | 40.1 | 51.9658 | 54.46 | 43.92 | 15.25 | — | — | 4 |
| 6 | Grok 4.5 | xAI | 49.1222 | 60.34 | 30.52 | 50.8918 | 62.08 | 33.17 | 100.07 | — | — | 3 |
| 7 | GLM 5.3 | Z.ai | 48.7517 | 60 | 27.86 | 55.814 | 71.7 | 29.59 | 173.93 | — | — | 3 |
| 8 | Kimi K3 | Moonshot AI | 48.1912 | 50.59 | 41.99 | 52.1741 | 58.94 | 41.32 | 208.66 | — | — | 3 |
| 9 | Opus 4.6 | Anthropic | 47.9068 | 61.77 | 27.86 | 57.2601 | 78.77 | 32.96 | 311.07 | — | — | 3 |
| 10 | DS-V4-Flash-0731 | DeepSeek | 47.3797 | 63.52 | 27.73 | 48.4328 | 74.24 | 29.96 | 56.46 | — | — | 9 |
| 11 | GLM 5.2 | Z.ai | 44.2628 | 48.2 | 35.12 | 43.3908 | 51.18 | 30.08 | 234.9 | — | — | 3 |
| 12 | GLM 5.3 Flash | Z.ai | 37.9708 | 52.98 | 17.8 | 40.3638 | 60.42 | 17.34 | 36.84 | — | — | 4 |
| 13 | Muse-Glimmer-30B | Meta | 25.8485 | 33.56 | 13.58 | 29.8485 | 43.7 | 14.08 | 52.62 | — | — | 9 |
| 14 | Nemotron 3 Ultra | NVIDIA | 18.7563 | 22.55 | 9.18 | 20.9019 | 24.25 | 11.94 | 330.12 | — | — | 3 |

### CyberGym

Domain: Security. Status: published.

Level 1 of CyberGym asks a model to reproduce 1,507 real C/C++ vulnerabilities in their upstream projects: given the repository and the advisory, produce an input that triggers the bug. A task scores only when the crash actually reproduces, so there is no partial credit for a plausible-looking attempt.

| Rank | Model | Provider | Score (pass@3) (percent) | Cost / task (USD) | Duration / task (seconds) | Standard error (percentage points) |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | DeepSeek V4.1 Flash | DeepSeek | 94.82 | 6.53 | 1353 | — |
| 2 | GLM 5.3 | Z.ai | 92.77 | 5.69 | 1283 | — |
| 3 | DeepSeek V4 Pro 0813 | DeepSeek | 92.1 | 8.83 | 1645 | — |
| 4 | Kimi K3 | Moonshot AI | 90.18 | 4.9 | 1657 | — |
| — | GPT-5.6 Sol* | OpenAI · rejected tasks | — | — | — | — |
| — | Claude Opus 5* | Anthropic · rejected tasks | — | — | — | — |

Note: GPT-5.6 Sol and Claude Opus 5 rejected the CyberGym task requests; no tasks were scored.

### LAB (Harvey)

Domain: Legal. Status: published. Tasks: 120.

Harvey's Legal Agent Benchmark hands a model the work a first-year associate would get: a partner-style instruction, a client matter holding both relevant and peripheral files, and a deliverable someone has to review. Tasks span 24 transactional, advisory, regulatory, and litigation practice areas. Two independent judges grade expert-written rubric criteria; the headline AP score awards half credit when only one judge gives an all-pass result.

Source: https://www.harvey.ai

| Rank | Model | Provider | Score (all-pass) (percent) | Score (Strict: both judges all-pass) (percent) | Score (Crit: criteria pass rate) (percent) | Cost / task (USD) | Duration / task (seconds) | Standard error (percentage points) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Grok 4.6 | xAI | 22.08 | 15 | 88.07 | 4 | 2160 | — |
| 2 | GLM 5.3 | Z.ai | 14.72 | 9.72 | 93.62 | 4.29 | 1560 | — |
| 3 | Kimi K3 | Moonshot AI | 13.06 | 9.72 | 91.87 | 4.86 | 2340 | — |
| 4 | Claude Opus 5 | Anthropic | 7.36 | 6.11 | 68.08 | 17.88 | 2280 | — |
| 5 | DeepSeek V4.1 Flash | DeepSeek | 5.69 | 3.61 | 78.02 | 0.21 | 1500 | — |
| 6 | GPT-6 Astra | OpenAI | 5.14 | 2.5 | 90.45 | 28 | 600 | — |
| 7 | DeepSeek V4 Pro 0813 | DeepSeek | 3.61 | 1.94 | 86.49 | 0.93 | 1380 | — |
| 8 | GPT-5.6 Sol | OpenAI | 0.56 | 0 | 82.24 | 1.29 | 360 | — |

### LAB: Contracts (Harvey)

Domain: Legal. Status: published. Tasks: 50.

The contracts slice of Harvey's Legal Agent Benchmark, covering review, redlining, and issue-spotting on real transaction documents. Two independent judges grade expert-written rubric criteria; the headline AP score awards half credit when only one judge gives an all-pass result.

Source: https://www.harvey.ai

| Rank | Model | Provider | Score (all-pass) (percent) | Score (Strict: both judges all-pass) (percent) | Score (Crit: criteria pass rate) (percent) | Cost / task (USD) | Duration / task (seconds) | Standard error (percentage points) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Grok 4.6 | xAI | 10 | 4.67 | 81.47 | 2.19 | 2160 | — |
| 2 | Kimi K3 | Moonshot AI | 8.33 | 4 | 89.75 | 5.54 | 2100 | — |
| 3 | Gemini 3.8 Flash | Google | 6.33 | 3.33 | 74.42 | 0.91 | 360 | — |
| 4 | GLM 5.3 | Z.ai | 6 | 1.33 | 89.76 | 4.31 | 1560 | — |
| 5 | GPT-6 Astra | OpenAI | 5.33 | 1.33 | 89.48 | 28 | 480 | — |
| 6 | GPT-5.6 Sol | OpenAI | 2 | 1.33 | 85.13 | 1.16 | 240 | — |
| 7 | Claude Opus 5 | Anthropic | 2 | 0 | 35.02 | 8.08 | 1200 | — |
| 8 | DeepSeek V4 Pro 0813 | DeepSeek | 1 | 0 | 81.13 | 0.93 | 660 | — |
| 9 | DeepSeek V4.1 Flash | DeepSeek | 0.33 | 0 | 23.97 | 0.29 | 1080 | — |

### APEX Agents: Corporate Law (Mercor)

Domain: Legal. Status: published. Tasks: 68.

Multi-step corporate-law workflows that chain tool use, retrieval, and document drafting into a single assignment. Practicing corporate attorneys grade the output against the work product a firm would actually accept.

Source: https://www.mercor.com

| Rank | Model | Provider | Score (avg@4) (percent) | Pass@1 (percent) | Pass@1 95% CI (+/-) (percentage points) | Cost / task (USD) | Cost / attempt (USD) | Attempt budget (k) (count) | Attempts scored (count) | Attempts expected (count) | Average tokens / attempt (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra | OpenAI | 87.67 | 73.16 | 9.74 | 4.422 | 1.1054 | 4 | 272 | 272 | 358442 |
| 2 | Claude Opus 5 | Anthropic | 83.8 | 69.49 | 9.93 | 5.772 | 1.4483 | 4 | 271 | 272 | 1330916 |
| 3 | Gemini 3.8 Flash | Google | 82.71 | 63.81 | 10.63 | 2.386 | 0.5965 | 4 | 268 | 272 | 2484096 |
| 4 | Grok 4.6 | xAI | 82.21 | 65.07 | 9.93 | 2.098 | 0.5245 | 4 | 272 | 272 | 605475 |
| 5 | GLM 5.3 | Z.ai | 81.41 | 67.65 | 9.56 | 2.492 | 0.623 | 4 | 272 | 272 | 784585 |
| 6 | GPT-5.6 Sol | OpenAI | 80.61 | 62.5 | 10.11 | 11.996 | 2.999 | 4 | 272 | 272 | 2836700 |
| 7 | DeepSeek V4 Pro 0813 | DeepSeek | 79.28 | 61.76 | 9.93 | 3.599 | 0.8997 | 4 | 272 | 272 | 1433792 |
| 8 | DeepSeek V4.1 Flash | DeepSeek | 78.37 | 61.76 | 10.11 | 0.56 | 0.1399 | 4 | 272 | 272 | 3556765 |
| 9 | Kimi K3 | Moonshot AI | 76.84 | 62.13 | 10.11 | 3.147 | 0.7866 | 4 | 272 | 272 | 562174 |

### Redline Bench (Crosby)

Domain: Legal. Status: published. Tasks: 140.

RedlineBench seats a terminal agent as in-house counsel and asks it to negotiate a contract across four sequential turns, editing the document in place with tracked changes and comments. The counterparty's instructions and the grading rubric stay hidden from the agent throughout. Scoring is criterion-by-criterion across twelve equally weighted scenario-and-turn cells.

Source: https://www.crosby.ai

| Rank | Model | Provider | Score (avg@3) (percent) | Cost / task (USD) | Duration / task (seconds) | Tool calls / task (count) | Uncached input tokens / task (tokens) | Input tokens (total) (tokens) | Cached tokens (total) (tokens) | Output tokens (total) (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra (max) | OpenAI | 62.8421 | 9.2568 | 2286.42 | 64.32 | 147541.6667 | 3962725 | 3815183.3333 | 79324 |
| 2 | Grok 4.6 | xAI | 54.715 | 4.6886 | 3163.84 | 94.13 | — | 5222638 | 4539903 | 175537 |
| 3 | Claude Opus 5 | Anthropic | 50.6682 | 9.1063 | 2835.26 | 103.56 | — | 7409399 | 7180707 | 163475 |
| 4 | GPT-5.6 Sol (medium) | OpenAI | 50.2755 | 1.1412 | 467.66 | 55.47 | 73390.4949 | 1023497.902 | 950107.4071 | 18051.6857 |
| 5 | GLM-5.3 | Z.ai | 49.5649 | 2.7216 | 1903.73 | 117.44 | — | 7059329 | 6677617 | 102513 |
| 6 | Kimi K3 | Moonshot AI | 49.0439 | 3.9484 | 1981.5914 | 86.4381 | — | 5446045.4 | 5149960.755 | 101008.9976 |
| 7 | DeepSeek V4 Pro 0813 | DeepSeek | 46.9099 | 0.632 | 858.52 | 121.35 | — | 5392633 | 5250233 | 53791 |
| 8 | Gemini 3.8 Flash | Google | 44.4785 | 1.8293 | 720.56 | 95.7 | — | 13295929 | 12553216 | 88215 |

### Big Finance Benchmark (Rogo)

Domain: Finance. Status: published. Tasks: 139.

Rogo's Big Finance Benchmark evaluates financial-research agents on private, workflow-grounded tasks. Models use Rogo's official agent loop and grading methodology, and the headline measures final-answer accuracy across independent full runs.

Source: https://github.com/Rogo-Technologies/big-finance-benchmark

| Rank | Model | Provider | Score (avg@3) (percent) | Score (rubric avg@3) (percent) | Score (pass@3) (percent) | Score (pass^3) (percent) | Cost (USD / task) | Cost (median USD / run) | Runs (full) (count) | Prompt tokens / task (tokens) | Cached tokens / task (tokens) | Completion tokens / task (tokens) | Reasoning tokens / task (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra | OpenAI | 54.6 | 68.3 | 57.2 | 49.3 | 0.303 | 71.6 | 3 | 140300 | 119700 | 1200 | 300 |
| 2 | Claude Opus 5 | Anthropic | 49.2 | 66.3 | 53.2 | 43.2 | 0.501 | 259.54 | 3 | 342000 | 0 | 8100 | 5700 |
| 3 | GLM 5.3 | Z.ai | 45.8 | 65.9 | 52.5 | 38.1 | 0.286 | 70.31 | 3 | 439500 | 400700 | 56400 | 53000 |
| 4 | GPT-5.6 Sol | OpenAI | 44.2 | 60.8 | 53.6 | 34.1 | 0.226 | 44.36 | 3 | 176800 | 152200 | 2500 | 1300 |
| 5 | Kimi K3 | Moonshot AI | 41.3 | 62.3 | 51.4 | 31.9 | 0.22 | 48.27 | 3 | 265400 | 236300 | 6700 | 4100 |
| 6 | DeepSeek V4.1 Flash | DeepSeek | 40.6 | 60.4 | 48.6 | 32.6 | 0.107 | 20.67 | 3 | 654900 | 610200 | 27500 | 24400 |
| 7 | DeepSeek V4 Pro 0813 | DeepSeek | 39.3 | 58.1 | 47.5 | 31.7 | 0.141 | 32.39 | 3 | 444500 | 407900 | 18600 | 16200 |
| 8 | GLM 5.3 Flash | Z.ai | 37.2 | 58.9 | 46.8 | 27.3 | 0.085 | 14.66 | 4 | 261100 | 219400 | 15700 | 13100 |

### DuetBench–Diagnosis (Decagon)

Domain: Customer Support. Status: published.

Real Duet diagnosis conversations replayed from the first human turn. A blinded judge compares model responses head to head on task outcome, investigation, tool use, and communication; the four relative ratings are averaged into Diagnosis Elo.

Source: https://decagon.ai/blog/duetbench

| Rank | Model | Provider | Score (Diagnosis Elo) (Elo points) | Task Outcome Elo (Elo points) | Investigation Elo (Elo points) | Tool Use Elo (Elo points) | Communication Elo (Elo points) | Completion rate (percent) | Pairwise battles (count) | Distinct cases (count) | Staleness (days) | Cost / task (USD) | Duration / task (seconds) | Uncached input tokens / task (tokens) | Cached input tokens / task (tokens) | Output tokens / task (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 1034.9 | 1039.7 | 1041.7 | 1025.6 | 1032.4 | 82 | 508 | 100 | 0 | 3.8629 | 326.3 | 563079 | 1165193 | 18597 |
| 2 | GLM 5.3 | Z.ai | 1031.5 | 1035.5 | 1029.3 | 1026.2 | 1034.8 | 98 | 240 | 50 | 0 | 0.8997 | 83.2 | 477051 | 731938 | 9437 |
| 3 | Claude 5 Sonnet | Anthropic | 1022.1 | 1019.2 | 1022.3 | 1017.6 | 1029.6 | 95 | 425 | 100 | 0 | 0.9968 | 239.3 | 347636 | 734031 | 15471 |
| 4 | Kimi K3 | Moonshot AI | 1011 | 1006.3 | 1018.7 | 1019.6 | 999.5 | 96 | 324 | 90 | 0 | 1.2164 | 143.2 | 280570 | 643584 | 12106 |
| 5 | GLM 5.2 | Z.ai | 1009.3 | 1008.6 | 1020.6 | 1007.2 | 1000.9 | 96 | 314 | 79 | 0 | 0.7641 | 189.3 | 400548 | 824466 | 19989 |
| 6 | Claude Opus 4.6 | Anthropic | 997.9 | 1003.5 | 979.7 | 1015.4 | 992.9 | 97 | 120 | 30 | 0 | 3.2718 | 315.2 | 517532 | 717019 | 13027 |
| 7 | GPT-5.6 Sol | OpenAI | 990.3 | 989.3 | 988.9 | 992 | 991 | 87 | 445 | 100 | 0 | 2.1212 | 118.1 | 446818 | 465313 | 7392 |
| 8 | GPT-5.6 Terra | OpenAI | 951.9 | 946.4 | 943.2 | 954.9 | 963.1 | 95 | 175 | 50 | 16 | 0.5489 | 65.2 | 214734 | 295627 | 5024 |
| 9 | GPT-5.6 Luna | OpenAI | 951.1 | 951.4 | 955.4 | 941.4 | 955.9 | 79 | 195 | 59 | 16 | 0.1019 | 85.9 | 420012 | 425882 | 7820 |

### τ³-Banking (Sierra)

Domain: Customer Support. Status: published. Tasks: 97.

Policy-bound banking agents handle realistic customer requests with database tools and a simulated user. The board covers the 97-task banking_knowledge subset, with four official seeded trials per task and binary rewards for correct final state and required natural-language assertions.

Source: https://github.com/sierra-research/tau2-bench

| Rank | Model | Provider | Score (mean trial reward) (percent) | Cost (USD / trial) | Duration (seconds / trial) | Uncached input tokens / task (tokens) | Cached tokens (total) (tokens) | Output tokens (total) (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GLM-5.3 | Z.ai | 51.6323 | 0.4852 | 281.6 | 108334 | 1056419 | 13382 |
| 2 | Kimi K3 | Moonshot AI | 51.5464 | 1.0179 | 422.6 | 145093 | 1381795 | 11206 |
| 3 | Gemini 3.8 Flash | Google | 51.3746 | 0.5313 | 285.3 | 302342 | 2763672 | 25945 |
| 4 | Qwen 3.8 Max | Alibaba | 51.2887 | 0.8748 | 293.6 | 202606 | 1483137 | 16464 |
| 5 | Grok 4.6 | xAI | 50.6873 | 1.3241 | 397.9 | 174545 | 1905867 | 3675 |
| 6 | Claude Opus 5 | Anthropic | 49.0979 | 2.4747 | 536.9 | 96097 | 2126013 | 37248 |
| 7 | GPT-6 Astra | OpenAI | 44.8454 | 1.805 | 410.6 | 40253 | 620696 | 15636 |
| 8 | DeepSeek V4 Pro 0813 | DeepSeek | 42.6546 | 0.3186 | 506.4 | 92762 | 1289019 | 35220 |
| 9 | DeepSeek V4.1 Flash | DeepSeek | 35.7388 | 0.0455 | 330.7 | 78687 | 936985 | 32785 |
| 10 | GPT-5.6 Sol | OpenAI | 35.6529 | 1.1411 | 334 | 90188 | 1257041 | 13879 |
| 11 | Muse Spark 1.2 | Meta | 34.2784 | 0.2977 | 228.4 | 44440 | 1348467 | 9382 |

### τ³-Voice (Sierra)

Domain: Customer Support. Status: published.

Full-duplex voice agents evaluated on task completion and natural interaction across retail, airline, and telecom support.

Source: https://sierra.ai

| Rank | Model | Provider | Score (Pass@1) (percent) | Cost / task (USD) | Duration / task (seconds) |
| --- | --- | --- | --- | --- | --- |
| 1 | gpt-live-1 | OpenAI · backend: gpt-6 astra (medium) · v1.0 | 81.7193 | — | 2.538 |
| 2 | Pine Voice Preview | Pine AI · enabled · v1.0 | 75.3801 | — | 2.1367 |
| 3 | grok-voice-think-fast-1.0 | xAI · enabled · v1.0 | 67.3216 | — | 1.2137 |
| 4 | grok-voice-think-fast-2.0 | xAI · high · v1.0 | 62.5263 | — | 1.7084 |
| 5 | qwen3.5-omni-plus-realtime | Qwen · — · v1.0 | 53.6725 | — | 1.6983 |
| 6 | gemini-3.1-flash-live-preview-thinking-high | Google · high · v1.0 | 43.848 | — | 3.1502 |
| 7 | gpt-realtime-2 | OpenAI · xhigh · v1.0 | 42.4327 | — | 1.9783 |
| 8 | gpt-realtime-2 | OpenAI · minimal · v1.0 | 38.5497 | — | 1.4395 |
| 9 | grok-voice-fast-1.0 | xAI · — · v1.0 | 38.3333 | — | 1.1493 |

### FrontierSWE V2 (Proximal)

Domain: Software. Status: published. Tasks: 34.

Thirty-four software-engineering tasks pitched at the edge of what an expert human can do: writing a flight-sim renderer in OpenGL, porting Git to Zig, driving a racing bot from vision alone. Each model gets five trials per task and up to twenty hours per trial, and every trial earns a graded reward rather than a pass or a fail, so a run that gets most of the way there still scores.

Source: https://www.frontierswe.com/

| Rank | Model | Provider | Score (mean@5) (percent) | Cost / task (USD) | Cost / attempt (USD) | Duration (seconds / trial) | Attempt budget (k) (count) | Attempts scored (count) | Attempts expected (count) | Worst@5–best@5 spread (percentage points) | Average tokens / trial (tokens) | Average steps / trial (count) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra | OpenAI | 65.5 | 5148.25 | 1029.65 | 43560 | 5 | 170 | 170 | 8.9 | — | — |
| 2 | Claude Fable 5.1 | Anthropic | 56.29 | 692.75 | 138.55 | 41760 | 5 | 170 | 170 | 11.1 | 246000000 | 623 |
| 3 | Claude Opus 5 | Anthropic | 52.01 | 983.55 | 196.71 | 57960 | 5 | 170 | 170 | 10.8 | — | — |
| 4 | GPT-5.6 | OpenAI | 32.2 | 898.2 | 179.64 | 30960 | 5 | 170 | 170 | 11 | 182000000 | 533 |
| 5 | GLM-5.3 | Z.ai | 30.18 | 486.1 | 97.22 | 61200 | 5 | 170 | 170 | 11.5 | 333000000 | 808 |
| 6 | Grok 4.7 | xAI | 29.5 | 1593.85 | 318.77 | 43560 | 5 | 170 | 170 | 12.6 | — | — |
| 7 | Kimi K3 | Moonshot AI | 25.87 | 548.55 | 109.71 | 66240 | 5 | 170 | 170 | 11.8 | 304000000 | 763 |
| 8 | Grok 4.6 | xAI | 25.29 | 1217.15 | 243.43 | 49680 | 5 | 170 | 170 | 12.2 | 229000000 | 882 |
| 9 | Gemini 3.8 Flash | Google | 19.63 | 193.75 | 38.75 | 25920 | 5 | 170 | 170 | 11.1 | — | — |
| 10 | DeepSeek V4 Flash Vision Exp | DeepSeek | 14.8 | 42.85 | 8.57 | 53280 | 5 | 170 | 170 | 9.7 | 513000000 | 1210 |

### APEX-SWE (Mercor)

Domain: Software. Status: published. Tasks: 200.

Repository-scale engineering tasks graded against the patch a maintainer would ship.

Source: https://www.mercor.com/apex/apex-swe-leaderboard/

| Rank | Model | Provider | Score (avg@4) (percent) | Pass@1 (percent) | Pass@1 95% CI (+/-) (percentage points) | Cost / task (USD) | Cost / attempt (USD) | Attempt budget (k) (count) | Attempts scored (count) | Attempts expected (count) | Average tokens / attempt (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 63.75 | 63.75 | 6.25 | 21.931 | 5.4829 | 4 | 800 | 800 | 3182369 |
| 2 | Grok 4.6 | xAI | 57.75 | 57.75 | 6.12 | 8.386 | 2.0966 | 4 | 800 | 800 | 2509336 |
| 3 | DeepSeek V4.1 Flash | DeepSeek | 51.88 | 51.88 | 6.25 | 0.978 | 0.2444 | 4 | 800 | 800 | 2620581 |
| 4 | GPT-6 Astra | OpenAI | 50 | 50 | 6.44 | 54.889 | 13.7222 | 4 | 800 | 800 | 1179698 |
| 5 | Kimi K3 | Moonshot AI | 48 | 48 | 6.25 | 2.589 | 0.6472 | 4 | 800 | 800 | 815738 |
| 6 | DeepSeek V4 Pro 0813 | DeepSeek | 46.92 | 46.92 | 6.21 | 8.045 | 2.0137 | 4 | 799 | 800 | 2175569 |
| 7 | GLM 5.3 | Z.ai | 45.62 | 45.62 | 5.94 | 12.777 | 3.1943 | 4 | 800 | 800 | 2596909 |
| 8 | Gemini 3.8 Flash | Google | 36.62 | 36.62 | 6 | 3.774 | 0.9482 | 4 | 796 | 800 | 4365138 |

### DeepSWE v1.1

Domain: Software. Status: published. Tasks: 113.

Deep debugging tasks where the failure cause sits several layers from the symptom.

| Rank | Model | Provider | Score (avg@3) (percent) | Cost / task (USD) | Duration / task (seconds) | Uncached input tokens / task (tokens) | Input tokens (total) (tokens) | Cached tokens (total) (tokens) | Output tokens (total) (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 72.27 | 6.1774 | 1309 | 137428 | 7469471 | 7332043 | 66103 |
| 2 | GPT-6 Astra | OpenAI | 71.01 | 2.559 | 681 | 60107 | 881158 | 821051 | 19746 |
| 3 | Kimi K3 | Moonshot AI | 70.21 | 5.3971 | 2340 | 190672 | 11799728 | 11609056 | 89493 |
| 4 | GLM 5.3 | Z.ai | 68.1433 | 4.7258 | 2294 | 270251 | 15351177 | 15080926 | 96907 |
| 5 | DeepSeek v4.1 Flash | DeepSeek | 63.4233 | 0.9177 | 1317 | 2479261 | 13063811 | 10584550 | 91982 |
| 6 | GPT-5.6 Sol | OpenAI | 61.95 | 2.0106 | 545 | — | 1517186 | — | 20025 |
| 7 | Gemini 3.8 Flash | Google | 61.65 | 4.2422 | 1180 | 1254164 | 30839486 | 29585322 | 288712 |
| 8 | Grok 4.6 | xAI | 57.8167 | 6.0288 | 5240 | 653412 | 6378373 | 5724961 | 309918 |
| 9 | DeepSeek V4 Pro 0813 | DeepSeek | 45.13 | 1.1231 | 1264 | 243277 | 11359099 | 11115822 | 79019 |

### MacroscopeBench (Macroscope)

Domain: Software. Status: published. Tasks: 195.

MacroscopeBench evaluates a model’s ability to identify bugs in real code changes while avoiding incorrect bug reports. Each task reviews a commit from a public repository; the evaluation includes commits containing known, subsequently fixed bugs and control commits with no known bugs recorded in the dataset. LLM judges determine whether the model identified the known bugs and assess the validity of all reported findings. The performance of a model is scored based on the harmonic mean of the known bug recall and overall bug detection precision as evaluated by the LLM judge.

Source: https://macroscope.com/benchmark

| Rank | Model | Provider | Score (avg@3) (percent) | Recall (any-of-k) (percent) | Recall (percent) | Precision (percent) | Recall (Critical severity) (percent) | Recall (High severity) (percent) | Recall (Medium severity) (percent) | Recall (Low severity) (percent) | Signal-to-noise ratio (valid findings per noise finding) | Cost / task (USD) | Duration / task (seconds) | Latency p90 (seconds) | Latency maximum (seconds) | Attempt budget (k) (count) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra · max | OpenAI | 77.9611 | 74.3056 | 67.8241 | 91.6606 | 45.2381 | 70.4545 | 77.5758 | 56.9892 | 6.2565 | 7.8746 | 228.86 | 478.725 | 935.58 | 3 |
| 2 | GPT-5.6 Sol · high | OpenAI | 77.7682 | 80.5556 | 70.8333 | 86.2083 | 50 | 69.697 | 77.5758 | 69.8925 | 3.7326 | 3.9031 | 180.398 | 330.454 | 1505.101 | 3 |
| 3 | GLM 5.3 · max | Z.ai | 77.379 | 87.5 | 76.3889 | 78.3951 | 69.0476 | 75 | 83.0303 | 69.8925 | 2.2392 | 3.8585 | 750.345 | 1367.982 | 2157.653 | 3 |
| 4 | Grok 4.6 · xhigh | xAI | 75.8716 | 78.4722 | 69.6759 | 83.2766 | 40.4762 | 68.9394 | 84.8485 | 56.9892 | 3.6658 | 4.6818 | 880.604 | 2735.451 | 3578.916 | 3 |
| 5 | Grok 4.6 · high | xAI | 75.0868 | 81.25 | 69.9074 | 81.0951 | 50 | 68.9394 | 83.0303 | 56.9892 | 3.1675 | 3.5108 | 929.989 | 2577.037 | 3580.73 | 3 |
| 6 | Kimi K3 · max | Moonshot AI | 72.8593 | 77.0833 | 66.6667 | 80.3201 | 47.619 | 65.9091 | 78.7879 | 54.8387 | 2.1811 | 2.8225 | 493.238 | 1334.473 | 3407.996 | 3 |
| 7 | Grok 4.6 · medium | xAI | 72.8008 | 75.6944 | 67.1296 | 79.5187 | 50 | 65.9091 | 83.0303 | 48.3871 | 3.1124 | 2.5609 | 518.544 | 2088.625 | 3529.837 | 3 |
| 8 | DeepSeek V4.1 Flash · max | DeepSeek | 71.9851 | 90.9722 | 83.1019 | 63.4916 | 71.4286 | 90.1515 | 84.2424 | 76.3441 | 1.2748 | 0.7506 | 977.285 | 1818.966 | 2816.489 | 3 |
| 9 | Gemini 3.8 Flash · medium | Google | 70.7316 | 70.8333 | 64.1204 | 78.8628 | 38.0952 | 65.1515 | 73.9394 | 56.9892 | 2.3374 | 3.3842 | 663.127 | 1589.509 | 2430.274 | 3 |
| 10 | DeepSeek V4.1 Flash · low | DeepSeek | 70.721 | 86.8056 | 77.5463 | 65 | 69.0476 | 76.5152 | 81.8182 | 75.2688 | 1.3699 | 0.384 | 489.42 | 1009.848 | 2156.195 | 3 |
| 11 | DeepSeek V4.1 Flash · high | DeepSeek | 70.503 | 85.4167 | 77.5463 | 64.6326 | 64.2857 | 79.5455 | 82.4242 | 72.043 | 1.3287 | 0.3762 | 460.21 | 927.771 | 1659.534 | 3 |
| 12 | Claude Opus 5 · high | Anthropic | 69.8731 | 82.6389 | 76.6204 | 64.218 | 57.1429 | 74.2424 | 87.8788 | 68.8172 | 1.204 | 6.5013 | 257.922 | 504.313 | 724.465 | 3 |
| 13 | GLM 5.3 · high | Z.ai | 69.4068 | 77.0833 | 64.5833 | 75.0089 | 50 | 65.1515 | 72.1212 | 56.9892 | 1.9357 | 2.0174 | 328.949 | 787.528 | 1860.395 | 3 |
| 14 | Kimi K3 · high | Moonshot AI | 68.5648 | 71.5278 | 62.037 | 76.6279 | 47.619 | 62.1212 | 72.1212 | 50.5376 | 1.887 | 1.8952 | 260.902 | 669.326 | 2541.027 | 3 |
| 15 | GPT-5.6 Sol · low | OpenAI | 66.1098 | 61.8056 | 55.0926 | 82.6347 | 39.2857 | 57.9545 | 65.4545 | 43.5484 | 3.1667 | 0.8059 | 63.909 | 114.25 | 298.891 | 3 |
| 16 | Claude Opus 5 · medium | Anthropic | 65.4963 | 76.3889 | 68.75 | 62.5366 | 47.619 | 70.4545 | 76.3636 | 62.3656 | 1.1815 | 3.496 | 131.671 | 310.143 | 858.378 | 3 |
| 17 | Claude Opus 5 · low | Anthropic | 59.3759 | 68.0556 | 57.4074 | 61.4843 | 40.4762 | 59.0909 | 64.2424 | 50.5376 | 1.116 | 1.4036 | 56.006 | 137.597 | 282.155 | 3 |
| 18 | Grok 4.6 · low | xAI | 57.1025 | 59.7222 | 52.5463 | 62.5239 | 42.8571 | 51.5152 | 64.2424 | 37.6344 | 1.305 | 0.298 | 54.939 | 113.728 | 729.184 | 3 |
| 19 | GLM 5.3 · low | Z.ai | 54.5613 | 61.8056 | 50 | 60.0383 | 52.381 | 50 | 57.5758 | 35.4839 | 1.1677 | 0.408 | 82.365 | 299.573 | 1268.588 | 3 |
| 20 | Kimi K3 · low | Moonshot AI | 53.5218 | 55.5556 | 45.1389 | 65.7284 | 38.0952 | 39.3939 | 55.1515 | 38.7097 | 1.3661 | 0.482 | 63.179 | 179.884 | 484.449 | 3 |
| 21 | Gemini 3.8 Flash · low | Google | 51.3346 | 50.6944 | 44.4444 | 60.7532 | 16.6667 | 48.4848 | 53.9394 | 34.4086 | 1.1062 | 0.9011 | 158.429 | 566.957 | 1848.816 | 3 |

### ORCA Benchmark (Traversal)

Domain: Software. Status: published. Tasks: 324.

Root-cause analysis on live incidents, narrowing traces, logs, and metrics to the failing change.

Source: https://www.traversal.com/blog/orca-bench-how-ready-are-language-model-agents-for-oncall

| Rank | Model | Provider | Score (Hard RCA, avg@3) (percent) | Score (Medium RCA, avg@3) (percent) | Hallucination (incidents) (percent) | Cost / task (USD) | Duration / task (seconds) | Uncached input tokens / task (tokens) | Input tokens (total) (tokens) | Cached tokens (total) (tokens) | Output tokens (total) (tokens) | Reportable runs (count) | Canonical verdicts per run (count) | Tool calls / task (count) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra | OpenAI | 50.21 | 65.99 | 3.75 | 3.9989 | 1143.73 | 1129568 | 1129568 | 0 | 38916 | 3 | 324 | 43.09 |
| 2 | Claude Opus 5 | Anthropic | 44.3038 | 68.0272 | 6.9913 | 5.3304 | 2597.7973 | 155908.0082 | 4901945.3827 | 4746037.3745 | 79325.3405 | 3 | 324 | 91.677 |
| 3 | Grok 4.6 | xAI | 32.07 | 45.24 | 5.49 | 2.4091 | 1636.87 | 351519 | 2977180 | 2625660 | 65540 | 3 | 324 | 61.22 |
| 4 | Kimi K3 | Moonshot AI | 31.22 | 52.04 | 10.24 | 1.24 | 1345.23 | 76894 | 1666781 | 1589887 | 35525 | 3 | 324 | 73.716 |
| 5 | GLM-5.3 | Z.ai | 29.96 | 56.8 | 6.37 | 2.3116 | 2418.7693 | 239623.3405 | 5534843.4671 | 5295220.1265 | 136215.6481 | 3 | 324 | 204.3909 |
| 6 | Gemini 3.8 Flash | Google | 29.54 | 42.86 | 8.24 | 1.4031 | 1139.548 | 639980.0597 | 9160068.2459 | 8520088.1862 | 75756.3313 | 3 | 324 | 101.3477 |
| 6 | GPT-5.6 Sol | OpenAI | 29.54 | 46.94 | 6.62 | 1.965 | 1435.82 | 72036 | 1112523 | 1040486 | 59434 | 3 | 324 | 110.46 |
| 8 | DeepSeek V4 Pro 0813 | DeepSeek | 24.47 | 37.41 | 9.74 | 0.5578 | 1128.49 | 128943 | 3948541 | 3819598 | 55436 | 3 | 324 | 119.49 |
| 9 | DeepSeek V4.1 Flash | DeepSeek | 20.25 | 41.16 | 9.99 | 0.1233 | 915.24 | 160272 | 5207222 | 5046950 | 79906 | 3 | 324 | 120.88 |

### Bedside Bench (Doximity)

Domain: Healthcare. Status: published. Tasks: 500.

Five hundred physician-authored clinical cases spanning ten suites of medical reasoning and safety, from differential diagnosis and workup through escalation and medication safety. Each case is graded by an LLM judge against criteria the authoring physicians wrote. The headline score is the macro-average of the ten suite means, so no single suite can carry a model.

Source: https://www.doximity.com

| Rank | Model | Provider | Score (avg@3) (percent) | Cost / task (USD) | Duration / task (seconds) | Standard error (percentage points) | Prompt tokens / task (tokens) | Cached tokens / task (tokens) | Completion tokens / task (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Doximity Ask V7 | Doximity | 95.4 | — | — | — | — | — | — |
| 2 | Claude Opus 5 | Anthropic | 94.6 | 0.1188 | 64.6 | — | 262.7 | 0 | 4698.7 |
| 3 | GPT-6 Astra | OpenAI | 93.4 | 0.106 | 53.6 | — | 161.2 | 0 | 2086.9 |
| 4 | GPT-5.6 Sol | OpenAI | 91.1 | 0.0501 | 42 | — | 161.2 | 0 | 2471.2 |
| 5 | Kimi K3 | Moonshot AI | 89.7 | 0.0529 | 82.9 | — | 250.2 | 2.3 | 3479 |
| 6 | GLM 5.3 | Z.ai | 86.9 | 0.021 | 94.7 | — | 106.9 | 59.3 | 4728.1 |
| 7 | Gemini 3.8 Flash | Google | 86.5 | 0.0092 | 10 | — | 155.4 | 0 | 2412 |
| 8 | Grok 4.6 | xAI | 83.3 | 0.0119 | 38.9 | — | 257.7 | 533.6 | 1855.7 |
| 9 | DeepSeek V4.1 Flash | DeepSeek | 82.9 | 0.0046 | 102 | — | 157.9 | 20.6 | 6923 |
| 10 | DeepSeek V4 Pro 0813 | DeepSeek | 80.6 | 0.0157 | 62.9 | — | 243.3 | 1.2 | 3878 |

### APEX-1: General Practitioner (Mercor)

Domain: Healthcare. Status: published. Tasks: 100.

Primary-care consults graded by practicing GPs on diagnosis, workup, and safe escalation. Each case is scored against the plan a GP would defend to a colleague rather than against one reference answer.

Source: https://www.mercor.com

| Rank | Model | Provider | Score (avg@4) (percent) | Pass@1 (percent) | Pass@1 95% CI (+/-) (percentage points) | Cost / task (USD) | Cost / attempt (USD) | Attempt budget (k) (count) | Attempts scored (count) | Attempts expected (count) | Average tokens / attempt (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5 | Anthropic | 72.12 | 9.25 | 4.87 | 1.971 | 0.4928 | 4 | 400 | 400 | 52484 |
| 2 | Gemini 3.8 Flash | Google | 68.36 | 6.75 | 4.25 | 0.204 | 0.0509 | 4 | 400 | 400 | 33634 |
| 3 | Kimi K3 | Moonshot AI | 64.35 | 5.25 | 4 | 0.692 | 0.173 | 4 | 400 | 400 | 30950 |
| 4 | GPT-5.6 Sol | OpenAI | 63.75 | 3.75 | 3.5 | 5.725 | 1.4312 | 4 | 400 | 400 | 190070 |
| 5 | GLM 5.3 | Z.ai | 63.2 | 5.5 | 3.88 | 0.445 | 0.1121 | 4 | 397 | 400 | 41779 |
| 6 | GPT-6 Astra | OpenAI | 61.03 | 6 | 4.38 | 0.709 | 0.1773 | 4 | 400 | 400 | 25006 |
| 7 | Grok 4.6 | xAI | 59.38 | 3.5 | 2.88 | 0.235 | 0.0588 | 4 | 400 | 400 | 27644 |
| 8 | DeepSeek V4 Pro 0813 | DeepSeek | 56.53 | 4.25 | 3.25 | 0.316 | 0.079 | 4 | 400 | 400 | 35442 |
| 9 | DeepSeek V4.1 Flash | DeepSeek | 56.18 | 4 | 3.25 | 0.056 | 0.014 | 4 | 400 | 400 | 36603 |

### HealthBench Professional

Domain: Healthcare. Status: published. Tasks: 525.

The professional-level slice of OpenAI's HealthBench: 525 clinical consults answered in free text and graded against physician-written rubrics. Scoring is length-adjusted, so restating the question earns nothing, and the headline is the mean weighted rubric score across every complete run.

| Rank | Model | Provider | Score (avg@8) (percent) | Cost / task (USD) | Duration / task (seconds) | Standard error (percentage points) | Prompt tokens / task (tokens) | Cached tokens / task (tokens) | Completion tokens / task (tokens) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | GPT-6 Astra | OpenAI | 64.43 | 0.0533 | 26.5 | — | 560.3771 | 0 | 953.5282 |
| 2 | GPT-5.6 Sol | OpenAI | 58.83 | 0.0727 | 56.8 | — | 560.3771 | 0 | 3521.653 |
| 3 | Claude Opus 5 | Anthropic | 55.4807 | 0.1399 | 76.1 | — | 893.8095 | 0 | 5418.6744 |
| 4 | Grok 4.6 | xAI | 51.661 | 0.0153 | 44.6 | — | 1183.4564 | 0 | 2157.3525 |
| 5 | Gemini 3.8 Flash | Google | 49.8163 | 0.0083 | 8.6 | — | 603.3352 | 0 | 2095.4964 |
| 6 | Kimi K3 | Moonshot AI | 48.8752 | 0.053 | 67.3 | — | 672.1314 | 0 | 3398.9333 |
| 7 | GLM-5.3 | Z.ai | 48.7264 | 0.0311 | 97.4 | — | 582.1752 | 0 | 6885.3539 |
| 8 | DeepSeek V4.1 Flash | DeepSeek | 47.76 | 0.005 | — | — | 580.4057 | 0 | 7391.2517 |
| 9 | DeepSeek V4 Pro 0813 | DeepSeek | 38.9052 | 0.0202 | 63.7 | — | 646.4057 | 0 | 4885.876 |

### Genspark Slides Benchmark (Genspark)

Domain: Productivity. Status: published. Tasks: 200.

Eleven models generate finished decks from 200 real production tasks in the Genspark Slides agent harness. Genspark's internal grader reports an aggregate score plus a quality composite, task completion, content quality, visual design, process quality, a layout-defect penalty, the Tier A deck share, and USD per deck.

Source: https://www.genspark.ai/blog/gen-1-slides

| Rank | Model | Provider | Score (percent) | Score (quality composite) (percent) | Score (task completion) (percent) | Score (content quality) (percent) | Score (visual design) (percent) | Score (process quality) (percent) | Penalty (layout defects) (percent) | Share (Tier A decks) (percent) | Cost (USD / deck) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Claude Fable 5.1 | Anthropic | 83.59 | 74.29 | 87.8 | 81 | 68 | 76.9 | 7.26 | 50 | — |
| 2 | Gen-1 Slides | Genspark | 82.22 | 75.21 | 77.8 | 73.4 | 75.8 | 72.2 | 21.52 | 71.4 | 0.44 |
| 3 | Claude Opus 5 | Anthropic | 81.02 | 73.14 | 84.9 | 78.2 | 68.7 | 69.6 | 11.72 | 50.8 | 4.16 |
| 4 | Claude Fable 5 | Anthropic | 79.92 | 70.53 | 87.2 | 81.1 | 61.4 | 77.1 | 12.13 | 40.7 | — |
| 5 | GPT-6 Astra | OpenAI | 77.95 | 65.96 | 85 | 78.3 | 54.5 | 81.8 | 3.47 | 21.1 | — |
| 6 | Kimi K3 | Moonshot AI | 72.56 | 66.21 | 84 | 78.4 | 55.9 | 74.9 | 21.67 | 34 | 2 |
| 7 | GPT-5.6 Sol | OpenAI | 67.02 | 60.58 | 82 | 74.2 | 48.2 | 74.9 | 21.31 | 22.3 | 2.01 |
| 8 | MiniMax M3 | MiniMax | 56.07 | 57.43 | 74 | 66.6 | 49.3 | 60.6 | 44.62 | 34.8 | 0.34 |
| 9 | GPT-5.6 Luna | OpenAI | 54.06 | 52.21 | 74.9 | 65.6 | 39.7 | 65.8 | 33.26 | 18.4 | — |
| 10 | Gemini 3.8 Flash | Google | 53.39 | 52.98 | 79.4 | 74.9 | 35.8 | 66.9 | 38.92 | 6.9 | — |
| 11 | GPT-5.6 Terra | OpenAI | 48.79 | 47.8 | 71.1 | 64.4 | 33.5 | 63.1 | 32.99 | 6.8 | — |
