Checkpoints and recovery
Save for the right purpose
| SDK operation | Purpose |
|---|---|
training_client.save_state(name=...).result() | Save training state for continued training |
training_client.load_state_with_optimizer(path).result() | Restore a training checkpoint including optimizer state |
training_client.save_weights_for_sampler(name=...).result() | Export weights for a sampling client; requires supported server configuration |
Wait for each future to complete before using its returned path. A sampler export does not replace the optimizer state needed to resume training.
Find the saved path
The SFT example writes latest-checkpoint.json to its output directory with path, step, model and dataset. This file is a locator; it is not the weight files. Keep it with your run records and use its exact returned path, without inventing a URI from a job ID.
With --verify-download, the example finishes the session, downloads durable files to verified-checkpoint/, verifies sizes and SHA-256, and writes durable-verification.json. Saving a locator, seeing a list entry or getting a successful script exit alone does not verify a complete local checkpoint download.
Resume explicitly
Use a checkpoint containing optimizer state, the same model and dataset, and the completed step from your run record:
uv run --locked python examples/train_sft_qwen_tulu3.py \
--resume-from CHECKPOINT_PATH --start-step 1 --max-steps 2 \
--batch-size 1 --save-every 1 --skip-sample \
--output-dir outputs/sft-resumed-run
Replace CHECKPOINT_PATH with the saved path and configure authentication as in the quickstart. --max-steps is the total target step count, so it must exceed --start-step. The example restores optimizer state and skips start_step × batch_size datums in the data stream. Preserve dataset ordering and batch size when recovering; a checkpoint does not freeze the external dataset stream for you.
Recovery may repeat work after the last checkpoint. Do not assume every gradient update executes exactly once. Checkpoint access remains owner-scoped; not every historical or cross-session checkpoint is accepted by the configured backend. If loading is rejected, check ownership and backend support with the administrator rather than changing the locator to bypass checks.
Sample from saved weights
When the administrator has enabled rollout in the policy, the example's --sample exports sampler weights, creates a sampling client from the returned model_path, calls sample(...).result() and writes sample.json. An actor-only policy does not become sampling-enabled when you add the flag.
Query registered checkpoints
Use the released CLI:
el train checkpoint list --session-id SESSION_ID --limit 25
The CLI lists registered metadata; it does not download weights. Empty results may mean the checkpoint has not been published or registered. Download access uses authorized signed URLs; it does not require giving users R2 credentials or sending a platform KEY to the storage URL.