SCOUT: On the Off-Policy Teacher in On-Policy Distillation
Student-COnditioned Updates of the Teacher
In on-policy distillation (OPD) the student writes the response and the teacher supervises every token of it. The teacher must continue text it would rarely write itself, and it does so less accurately the longer the student's prefix gets. Earlier fixes decide when to trust a frozen teacher. SCOUT trains the teacher instead: the student keeps standard OPD, and every few steps the teacher learns, with outcome-reward RL, to continue from the student's own prefixes. The teacher gets better at exactly the states it has to supervise, and the student improves across teachers, model families and domains.
On-policy for the student is off-policy for the teacher
In on-policy distillation (OPD) the student writes the response and a stronger teacher scores every token of it, so the student learns from its own samples.
A teacher is trained to continue its own reasoning. In OPD it has to continue the student's instead: every token it scores comes after text the student wrote. Small differences and mistakes in the student's steps pile up, so the further into a response, the further the text drifts from anything the teacher would have written. We call this the off-policy teacher problem.
You can see it directly. Take a student response, cut it partway, and ask the teacher to finish the solution. Drag the slider to give the teacher a longer piece of the student's response.
Can the teacher finish the student's solution?
Accuracy when the teacher continues from the cut
How unsure is the teacher along the response?
Next-token entropy of the teacher (higher = less sure)
Is it just running out of tokens?
No. The teacher's continuation uses less than half of its budget on average.
Accuracy falls from about 47% to 30% as the prefix grows, and the teacher stays unsure on student text while growing more certain on its own. So the teacher is least reliable deep into the response, which is where OPD still asks it to supervise every token.
Earlier fixes accept this and keep the teacher fixed. They distil only the early part of the response (ESR), supervise only where teacher and student are compatible (Prune-OPD), or let the teacher step in during the rollout (Relay-OPD). We ask the complementary question: can we train the teacher itself to supervise student-generated prefixes better?
Train the teacher to continue the student's prefixes
SCOUT keeps the student's OPD update unchanged and adds one step: every few training steps, the teacher practises finishing the student's own partial responses, and is rewarded when it reaches the correct answer.
Prefix curriculum. Long student prefixes are the hardest for the teacher, so the cut starts early in the student's response and moves steadily later as training goes on. Drag the slider to move it.
The best distilled student in every setting
Four settings: a Qwen3 pair on math, the same pair on code, a larger RL-trained 8B teacher, and another model family (Skywork teaching DeepSeek). We compare with OPD, the three fixed-teacher variants above, and GRPO, which trains the student with RL on correct answers and no teacher.
Each method is trained 3 times unless marked. Scores are Avg@k: accuracy averaged over k sampled answers per question.
SCOUT generalizes to different settings
Starting from 4B → 1.7B math (+2.2 over OPD), we change one thing at a time.
-
+2.6vs OPDA different, larger teacher. With Qwen3-8B-DAPO as the teacher, SCOUT is best or tied-best among OPD methods on all six math benchmarks, as it is with the 4B one.
-
+3.1vs OPDCode. Same 4B → 1.7B pair, graded by unit tests. Relay-OPD falls to 12.1: its code often contains syntax errors. GRPO, which needs no teacher, is still stronger on code (62.2).
-
+1.2vs OPDAnother model family. Skywork-OR1-Math-7B teaching DeepSeek-R1-Distill-Qwen-1.5B. Here GRPO collapses during training and Prune-OPD falls well below OPD.
Where the gains come from
We answer five questions about SCOUT:
- Does the teacher get better at continuing student prefixes?
- Does the gain come from the student prefixes, or just from training the teacher more?
- How often does the teacher need to update?
- Does training OPD longer close the gap?
- Does SCOUT combine with other OPD improvements?
RQ1Does the teacher get better at continuing student prefixes?
Yes. We use checkpoints from the 4B → 1.7B math run and ask the teacher to finish student responses cut at different points on AIME 2025.
Matched pairs (Figure 2): student and teacher come from the same training step. As training goes on, the curves shift up, and later pairs stay accurate on longer student prefixes. But a later student also writes better prefixes, so this mixes up the two models. Fixed student separates them: every prefix comes from the student before training, and only the teacher changes. The trained teacher is more accurate at every prefix length, so the gain is in the teacher itself.
Uncertainty (Figure 3): before training, the teacher is much less sure on student text than on its own. After SCOUT, its uncertainty on student text drops over the later part of the response, and the gap to its own text narrows.
RQ2Student prefixes, or just more teacher training?
The student prefixes. SCOUT does two things: it trains the teacher more, and it trains it on student prefixes. To separate them, OPD + Teacher GRPO gives the teacher the same RL updates, but its rollouts start from the problem itself instead of from a student prefix.
Extra teacher training alone explains only part of the gain. It is slightly below OPD on 4B math and gives smaller gains on 4B code and 8B math. SCOUT is the strongest in all three settings: the rest comes from training the teacher on the text the student actually writes.
RQ3How often does the teacher need to update?
Not every step. In the 8B → 1.7B setting, updating the teacher every 1, 5 or 10 steps works about equally well; sparser updates (every 15 or 20 steps) fall back to about OPD's level. We use every 10 steps.
Q4Does training OPD longer close the gap?
Not on the same data. SCOUT spends extra compute on teacher rollouts and teacher updates. To see whether OPD can make up the difference with that budget, we train OPD for three epochs and SCOUT for one, on the same data, so both use similar training compute. OPD improves at first, then stays around 49–50 through its second and third epochs. SCOUT keeps improving and ends higher (51.7 vs 50.4).
Two limits apply. The extra OPD epochs reuse the same data, so fresh data could still help OPD. And at matched compute, SCOUT still takes more wall-clock time and memory.
Q5Does SCOUT combine with other OPD improvements?
Yes. SCOUT improves the teacher; loss-level methods change how the student uses the teacher's signal. We combine SCOUT with On-Policy Trust Region (OPTR), the first component of Trust Region On-Policy Distillation (TrOPD). OPTR already beats OPD, and adding SCOUT improves it by a further +2.5 on math and +1.2 on code. On math the two gains add up; on code they partly overlap.
Math · Mean of 6 benchmarks
Code · Mean of LCB v5, HumanEval+, MBPP
BibTeX
@article{huang2026scout,
title = {{SCOUT}: On the Off-Policy Teacher in On-Policy Distillation},
author = {Huang, Langlin and Liu, Hao and Goswami, Mononito and Li, Xinyu and
Jana, Prithwish and Kanakaris, Nikos and Bl{\"o}baum, Patrick and Jain, Purak},
journal = {arXiv preprint},
year = {2026}
}