Hosted SFT is in closed beta — Prime enables it per account. Contact Prime support if dispatch is refused.
prime-rl supervised fine-tuning on Prime’s hosted GPU clusters: create a volume, copy your dataset onto it over SSH, dispatch. No kubectl, no cluster credentials — the trainer reads the dataset from the volume.
This walkthrough fine-tunes a small Qwen3 model to reverse text (reverse-text) on one GPU. For SFT on your own infrastructure, see Training.
Prerequisites
prime CLI v0.9.0 or later.
- A Prime account with hosted training access — Prime-side enablement; the platform deployment needs the hosted SFT path switched on.
- A model cached on the cluster that owns your volume — shared-LoRA models are not valid here. Check with:
- A Hugging Face dataset you can download with
hf download— withprompt/completionormessagescolumns (Training § Dataset Format).
1. Create a volume
Size for what you keep: datasets plus weights + optimizer state per checkpoint — the 0.6B smoke run wrote ~7 GB; a 100B-class
[ckpt] needs terabytes. Grow with prime volumes resize <name> --size <size>.2. Put the dataset on the volume
SSH into the volume and download the dataset onto it. The session runs next to the volume in the cluster, so the download uses the cluster’s network — not your laptop:/volume — the same path you see in the SSH session. Local files work too: prime volumes put <name> ./my-data /volume/datasets/my-data. In the config, data.name accepts either form: a path relative to the volume (datasets/my-data) or the absolute mount path (/volume/datasets/my-data) — the platform resolves both to the same location.
3. Write the config
Save this assft.toml — data.name points at the dataset on the volume (relative to the volume root; absolute /volume/... paths work too):
Hosted SFT is trainer-only:
[eval], [inference], and [weight_broadcast] blocks are rejected at dispatch, and output_dir is injected by the platform — omit both. The walkthrough model is a cached debug sibling already fine-tuned on this dataset — treat the loss curve as a smoke test of the hosted path.4. Dispatch
<runId>) below. The run moves PENDING → CREATING → RUNNING → COMPLETED — expect roughly three minutes on one H200 for this walkthrough.
5. Watch it run
datasets/ tree stay on the volume.
Fetch your outputs any time:
0 tokens / $0.00 for SFT — a reporting quirk, not free usage.)
Troubleshooting
[eval] / [inference] blocks rejected
[eval] / [inference] blocks rejected
Hosted SFT is trainer-only —
[weight_broadcast] is rejected for the same reason. Remove the blocks, or run locally with uv run sft.Model not available on this cluster
Model not available on this cluster
Hosted SFT boots from the cluster model cache. Check
prime train models --fft-only, or ask Prime support to cache the model.data.name requires --volume
data.name requires --volume
Real data needs a volume.
data.type = "fake" is the one no-volume path (synthetic data, connectivity smoke test).Dataset missing on the volume
Dataset missing on the volume
The trainer reads
data.name from the volume mount — a path that is empty or missing fails at startup. Re-run step 2 and check the name matches: the directory you create under /volume/datasets/<name> in the SSH session is exactly the datasets/<name> you point data.name at (relative paths resolve under /volume, so datasets/<name> and /volume/datasets/<name> are the same location).Volume ran out of space during download
Volume ran out of space during download
Inside the SSH session,
hf download fails when the volume is full. Exit, grow the volume, and re-run the download:Limitations
- Requires CLI v0.9.0+ — older versions misroute SFT configs against the RL schema.
- SFT-shaped datasets —
prompt/completionormessagescolumns. - Cluster-cached models only —
prime train models --fft-only; support can add others. - Trainer-only — no online evals;
[eval]/[inference]/[weight_broadcast]rejected. - Checkpoints live on the volume —
prime train checkpointsdoesn’t list SFT ones; useprime volumes get/ssh.
Scaling up
The same flow carries to multi-node — switch[deployment] to multi_node and add the parallelism knobs. Save as glm53-sft.toml:
This config targets 8×8 H200 (64 GPUs) — 4 steps take ~5.5 min on that hardware. If you adapt it: remove
force_balanced_routing — it forces round-robin expert routing and is debug-only; with warmup_steps = 50 a 4-step run never leaves warmup, so losses are warmup-dominated rather than convergence; add a [ckpt] block if you want checkpoints (this one omits it, so none are saved); with this linear schedule the LR ramps to its peak at step 50 — expect a transient loss spike near the peak before it settles (9.42 at step 55 in a 60-step run). See Scaling and Advanced for the knobs.Full Fine-Tuning
RL-style full-parameter training on a dedicated cluster — same CLI, same dashboard.
prime-rl Training
The underlying SFT trainer: config schema, dataset format, and local runs.