flash train. Flash runs the
job on managed infrastructure, supervises it, and streams checkpoints and logs
back to you. The config and run lifecycle follow; the
configuration reference lists every field.
Pick a base model
Flash trains a LoRA adapter on top of a supported base model. List base model ids and their parameter sizes (Supported models covers algorithms, reasoning, and pricing):model_revision pins all student-model loads to one Hugging Face revision. An
empty or omitted value uses the model’s default revision. seed controls the
run’s deterministic training order and defaults to 42.
Choose a training algorithm
sft: supervised fine-tuning, when you already have the answers. The model imitates the prompt/answer pairs in your environment’s dataset.grpo: reinforcement learning, when there’s no fixed answer to copy. Your environment’s reward scores each completion.opd: on-policy distillation, when a stronger model already does the task. A managed teacher, GLM 5.2 by default or another selected withteacher_model, grades your model’s own completions token by token. Training pulls your model toward that teacher without using answers or a reward as the training signal. Warm-start from an SFT adapter (init_from_adapter) for best results; a cold OPD run tends to underperform SFT. OPD supports single- and multi-turn environments, but not tool-calling ones. Teacher access is managed and teacher tokens are not billed to you.
algorithm defaults to sft when omitted. See
how Flash works for the difference.
Check the teacher before OPD
OPD pulls your model toward the teacher, so the teacher’s competence on your task roughly sets the ceiling. Treat a teacher evaluation as a precondition for the run. Flash does not automate this comparison. Use your evaluation harness, or have the coding agent follow the scaffoldedTRAINING.md, to call the teacher on held-out
prompts and score its transcripts with the same scorer you use for your model.
- Keep a held-out split outside training, then run both the selected teacher and your current model on it.
- Compare their scores, then read several trajectories from both. A score alone can hide a teacher that reaches the right answer for the wrong reason.
- Confirm that the teacher clearly wins and that your environment rewards the behavior you want to transfer.
The environment score does not drive OPD’s token-level updates. It tells you
whether the behavior you care about improved, so a broken evaluation leaves you
unable to judge the run.
Bound reasoning before long rollouts
If the teacher over-reasons in the held-out evaluation or OPD rollouts hit the length limit, put a hard, specific budget in the environment’s system prompt:Anatomy of a config
structured_outputs. See the Structured outputs
guide.
Control the update horizon and checkpoints
For SFT, GRPO, and OPD, a positivetrain.max_steps is the exact number of
optimizer updates. If it is absent or non-positive, Flash derives the update
count from epochs, retained examples, batch size, and the recipe.
Use train.save_at_steps when specific checkpoint boundaries matter:
max_steps, and cannot contain a step beyond it. It suppresses save_every,
and every requested save is mandatory. Without save_at_steps, periodic
save_every uploads block training while attempted but remain best-effort after
bounded retries. Use save_at_steps for required durable boundaries.
For SFT, max_examples = N retains the first N rows in file order before the
seeded shuffle. SFT also honors max_context_tokens for the rendered sequence.
Warm-start safely
init_from_adapter continues a source adapter in a new GRPO or OPD run. The
source can be a finished SFT run or a saved checkpoint. Its LoRA rank and alpha
are authoritative, so omit both lora_rank and
lora_alpha. The server resolves them during --dry-run and real submit; local
--cost may warn and estimate with defaults.
If you set model_revision, the source and target revisions must match exactly.
Use server-side --dry-run to resolve and validate the source.
Infrastructure is managed
You choose the model, algorithm, environment, and training settings. Flash selects suitable infrastructure by default. For controlled runs,[gpu] can pin
a provider and an active validated GPU class from flash gpus; those constraints
also apply to retries and cost estimates. Artifact storage remains managed.
Validate before you submit
--dry-run is authenticated and server-side. It runs submit-time preflights,
including warm-start resolution, creates a run with state=dry_run, and reports
client/server [train] schema differences. It allocates no GPU and charges
nothing:
--cost separately for the local estimate.
Submit the run
flash train follows the logs until the run finishes. Useful flags:
Cost and billing
Before submitting, Flash checks your org’s prepaid balance against the pre-flight estimate. A successful run bills at the quoted Flash cost. A run cancelled after training starts is repriced to the steps it reached. Setup time is reported but not billed. Preview the estimate before you submit:Monitor a run
Ctrl-C during flash train detaches you. The run keeps going on Freesolo.
After training
When a run reachesdone, serve it with
flash deploy and talk to it with flash chat.
Deploy or
export a run
you want to keep. Freesolo’s managed storage garbage-collects a run’s
checkpoints and final adapter about 7 days after its last activity when
the run was never deployed. A deployed run, or one that another run
warm-started from, is kept. Exporting copies the adapter into a HuggingFace
repo you own.
Next steps
Deploy & chat
Serve a finished run and talk to it.
Configuration reference
Every field you can set in a training config.