Skip to main content

SFT quickstart

This example fine-tunes Qwen/Qwen3.8-27B using allenai/tulu-3-sft-mixture. It runs real training and can incur GPU costs. Use a workspace with an administrator-approved model policy and budget.

Prepare the example​

Use Python 3.11+ and uv. The runnable example lives in EnvPlatform and needs its sft dependencies, model tokenizer access and dataset access:

git clone https://github.com/EnvLoop/EnvPlatform.git
cd EnvPlatform
uv sync --locked --extra sft
uv run --locked python examples/train_sft_qwen_tulu3.py --help

If you already have a checkout, use it. Installing the query-only envloop-cli package does not install this example or its dependencies.

Authenticate​

Create a platform machine KEY in Dashboard → API Keys for your authorized workspace. Follow Installation to save it in a private file outside the repository with mode 0600.

export ENVLOOP_BASE_URL=https://api.envloop.ai
export ENVLOOP_WORKSPACE_ID=YOUR_WORKSPACE_ID
export TINKER_CREDENTIAL_CMD='cat /absolute/private/platform.key'
unset TINKER_API_KEY

Replace the workspace and absolute file path. TINKER_CREDENTIAL_CMD lets the SDK read the platform KEY to exchange and refresh short-lived credentials. It is different from the CLI's ENVLOOP_API_KEY. Do not use a provider Tinker KEY or leave TINKER_API_KEY set, even to an empty string.

Use the API origin without a path. The example appends /tinker and sends X-Workspace-ID; you do not connect directly to GPU hosts or configure cloud credentials.

Run one step​

Choose a fresh output directory for each run:

uv run --locked python examples/train_sft_qwen_tulu3.py \
--max-steps 1 --batch-size 1 --max-samples 10 --save-every 1 \
--skip-sample --verify-download \
--output-dir outputs/sft-first-run

The script validates the first training datum before creating a paid session, creates a LoRA training client, performs forward_backward(...).result() and optim_step(...).result(), saves a checkpoint, then finishes with close(...).result(). With --verify-download, it verifies durable checkpoint files after finish. Provisioning, model loading and compilation can take time; wait for operation results rather than treating a session ID as readiness.

Text-only system/user/assistant messages are supported; loss applies to assistant labels. The example does not silently truncate data. LoRA rank and backend resources come from the server policy.

Adjust the run​

OptionDefault / behavior
--modelQwen/Qwen3.8-27B; must match an approved policy
--datasetallenai/tulu-3-sft-mixture, streaming train split
--max-steps / --batch-size10 / 2
--lr or --learning-rate2e-4
--save-every10; the final completed step is also saved
--max-samplesOptional dataset row limit, not generated sample count
--sample / --skip-sampleSampling disabled by default; enabling it requires a rollout-enabled server policy
--verify-downloadFinish, download and verify checkpoint file sizes and SHA-256

A larger batch, context or enabled sampler needs its own resource validation. The recorded two-GPU actor-only smoke does not validate sampling.

Use envloop training imports​

For your own training script, install the envloop source checkout with its training extra:

# In a separate envloop repository checkout
git clone https://github.com/EnvLoop/envloop.git
cd envloop
uv sync --locked --extra train
from envloop import training
from envloop.training import types

These are direct Tinker 0.30.2 public aliases. Use training.ServiceClient for SDK operations and import types as shown; deep imports such as envloop.training.types are not provided. The envloop SFT example additionally requires EnvPlatform helpers and SFT dependencies; see its source setup.

Next: save and recover checkpoints, or inspect the job.