- In training: use a JSON schema, regex, fixed choice set, or any-JSON constraint so your reward measures the answer’s content instead of malformed wrappers.
- At serving: the training constraint becomes the adapter default. The
OpenAI-standard per-request override supports
text,json_object, andjson_schema. Give the request enough output tokens and checkfinish_reasonfor truncation.
In training
structured_outputs in the [train] table constrains every GRPO and OPD rollout.
It applies only to the rollout algorithms — SFT never samples, so it rejects the
key at submit. When rollouts can’t drift off-format, your score_response parses
the field it cares about directly and rewards the content, which makes the reward
easier to write and the signal cleaner.
Write a constraint
Set exactly one constraint under[train]. A JSON schema is the common case;
write it as a JSON string or an inline TOML table.
A bare schema table with no constraint key is read as a JSON schema — its
type/properties keys are schema vocabulary, not Flash’s. Setting more than one
constraint, or the unsupported grammar / structural_tag forms, is rejected at
submit. To turn guided decoding off, omit the key or set false, "", or
"none".
Options
These optional keys tune the constraint. Set them alongside a constraint (an option on its own is rejected):Thinking mode
Structured outputs work withthinking = true. For a single-turn rollout and
the first assistant turn of a multi-turn rollout, reasoning remains free-form and
the grammar begins after the model closes </think>. The final answer must
satisfy the configured constraint.
A later multi-turn prompt may already contain an earlier </think> boundary, so
its grammar can apply from the first new token. Test later-turn behavior when a
multi-turn environment combines thinking and structured outputs.
Keep the reasoning and final answer within max_completion_tokens. An unclosed
or truncated thinking block never reaches the constrained answer and is invalid.
Score a completed constrained answer
For a non-truncated rollout that reaches the constrained final answer, the wrapper is guaranteed. Still return zero if reasoning or generation ends before that answer is complete:At serving
A deployed adapter or a base model accepts the OpenAI-standardresponse_format.
A training structured_outputs constraint becomes the deployed adapter’s
default. A request-level response_format overrides that default without a
redeploy:
{"type": "json_object"} forces any valid JSON with no fixed schema, and
{"type": "text"} leaves output unconstrained. With thinking enabled, reasoning
remains free-form and the requested grammar begins after </think>. Give each
request enough output tokens and check finish_reason before parsing the answer.
Every real adapter deployment runs a bounded smoke when a default constraint is
present. Deployment fails before alias activation if output is invalid or leaves
thinking unclosed or truncated. See
Deploy & chat.
Next steps
Configuration reference
The
structured_outputs field in the training config.Deploy & chat
Constrain a served model’s responses per request.