This page lists supported models and per-token pricing for the legacy shared LoRA service. Full fine-tuning (FFT) is currently in closed beta.
Shared Hosted Training for LoRA runs will stop accepting new runs on October 8, 2026. Hosted Training is transitioning to dedicated runs. Existing LoRA adapters will remain downloadable and deployable until further notice.
Available Models
Prices are per million tokens, billed separately for input, output, and training.
Trace storage charges
Traces uploaded to Prime Traces incur separate ingestion and storage charges. They are not included in the model token rates above. The workspace’s Usage Limits contains the upload control.
The November 6 legacy sample storage retirement is separate from the October 8 shared LoRA service shutdown.
Choosing a Model
For Validation and Debugging
Start with a small, fast model to verify your environment and config work correctly before committing compute to a larger run.
Recommended: Qwen/Qwen3.5-0.8B or meta-llama/Llama-3.2-1B-Instruct
For Experimentation
MoE models with small active parameter counts give strong performance per token.
Recommended: Qwen/Qwen3.5-35B-A3B
For Production Training
For serious training runs where you want the strongest results available, use a flagship MoE model.
Recommended: Qwen/Qwen3.6-35B-A3B
Thinking Mode
Qwen3.5 and Nemotron models support a thinking mode that produces extended chain-of-thought reasoning before the final answer. Toggle it via [sampling].enable_thinking in your config. Thinking mode tends to help on tasks that benefit from multi-step reasoning (math, code, logic) at the cost of longer outputs.
Checking Available Models
Always use the CLI to check the current list of supported models:
Add --output json to get the live pricing alongside the model list. This may differ from this page as models are being added regularly.