Base model parameter count | $ / 1M input tokens |
|---|---|
up to 150M | $0.008 |
150M - 350M | $0.016 |
Qwen3 8B | $0.1 |
Supervised and preference fine tuning is priced per 1M training tokens.
Base Model | LoRA SFT | LoRA DPO | Full Param SFT | Full Param DPO |
|---|---|---|---|---|
Models up to 16B parameters | $0.50 | $1.00 | $1.00 | $2.00 |
Models 16.1B - 80B | $3.00 | $6.00 | $6.00 | $12.00 |
Models 80B - 300B (e.g. Qwen3-235B, gpt-oss-120B) | $6.00 | $12.00 | $12.00 | $24.00 |
Models >300B (e.g. DeepSeek V3, Kimi K2) | $10.00 | $20.00 | $20.00 | $40.00 |
Attach to a shared, always-on trainer pool for LoRA training on the launch models. There's no provisioning and no idle cost. You pay only for the tokens you prefill, sample, and train.
Base Model | Context | Prefill / 1M | Cached Prefill / 1M | Sample / 1M | Train / 1M |
|---|---|---|---|---|---|
GLM 5.3 | 262K | $4.86 | $0.972 | $12.15 | $14.58 |
Qwen 3.8 27B | 128K | $1.86 | $0.372 | $5.595 | $4.103 |
Kimi K3 | 192K | $10.87 | $2.17 | $27.11 | $32.55 |
DeepSeek V4 Flash 0731 | 262K | $1.74 | $0.35 | $4.33 | $5.20 |
Muse Glimmer 30B | 128K | $1.96 | $0.39 | $4.88 | $5.86 |
Dedicated Training API jobs are priced per GPU hour. Please see the On-Demand Pricing section below for details on Dedicated Training API Pricing.
GPU Type | Price ($) per minute | Price ($) per hour |
|---|---|---|
H100 80 GB GPU | $0.134 | $8.00 |
H200 141 GB GPU | $0.134 | $8.00 |
B200 180 GB GPU | $0.217 | $13.00 |
B300 288 GB GPU | $0.250 | $15.00 |
GB300 288 GB GPU | $0.334 | $20.00 |