Skip to main content
The flash-example repository contains eight small, self-contained workflows across pure SFT, single-stage OPD, SFT-to-GRPO warm starts, and SFT-to-OPD warm starts. Each example packages its environment, reward-verified training rows, frozen 50-case held-out set, training config, and evaluation entrypoint. The results are task-specific evidence, not a general model ranking. Several students reach parity or near parity with GPT-5.5 under their exact boxed-answer, tool-use, or multi-turn contracts, and some score differences reflect strict protocol compliance rather than broad capability superiority. The numerical source of truth, including run ids, corrected-result history, footprint notes, and SFT-versus-RL ablations, is RESULTS.md.

Explore the flash-example repository

Browse the environments, training configs, frozen evaluations, and the task-specific numerical record in RESULTS.md.