

forward_backward for a built-in loss, or forward_backward_custom when the loss is yours. optim_step applies the update.save_weights_for_sampler and DCP snapshots, then promote the checkpoint you want to serve. inactivityTimeout, or disableable) to avoid runaway cost. Deployments aren't auto-cleaned; delete or scale them down.Training API Compute Options
| Serverless Compute (Self-Serve) | Dedicated Compute (Talk to us) | |
|---|---|---|
| Provisioning | Attach to an always-on shared pool | Trainer and deployment provisioned for your run |
| Billing | Per token, no idle GPU charge | Time-based on the GPUs you hold |
| Models | A curated list of popular models | Unlimited depth and breadth, up to the largest MoE models |
| Parameter mode | LoRA | LoRA and full-parameter |
| Throughput | Shared capacity and per-account rate limits | No contention, no rate limits, scale up as much as you need |

Three built-in methods. Which one fits depends on the data you already have.
Select from our model catalog; eligibility is set per model and per method.
OpenAI-compatible chat completion format, so existing OpenAI SFT datasets run without conversion.
firectl and the API, or hand the job to your coding agent.