Small-Model Distillation, Part 5: Training a 0.8B SQL Model with Its Own Feedback
TL;DR
This post tests on-policy self-distillation on two already-trained versions of Qwen3.5-0.8B. The student generates actions, and a frozen copy of its starting checkpoint supplies next-token probabilities for training. Checkpoint A had already learned from teacher examples and corrections to student mistakes. Part 4 then trained A with probabilities from a 35B teacher to produce Checkpoint B.
| Starting checkpoint | Training tested here | Result on 220 SQL tasks |
|---|---|---|
| A: 67 solved | Self-distillation using frozen A | 69 solved after 5 steps; 63 after 10 |
| B: 69 solved | Self-distillation using frozen B | No improvement; best result was 67 |
Self-distillation from A matched Part 4’s best larger-teacher result. Adding it after the larger-teacher stage, starting from B, did not improve that result. A separate variant gave the small teacher reference solutions or attempt feedback while training from B; it also scored below B. The evaluation scores guided candidate selection, so the gain from A needs confirmation on new tasks.
Where Part 4 left off
The SQL agent and its task carry over from the earlier posts. Given a question, the model inspects a SQLite database, runs queries, reads the results, and submits final SQL. A task counts as solved when that SQL passes hidden tests. The tools, action format, parser, turn limit, and scorer stayed fixed for this experiment.
Before Part 4, the 0.8B model had learned from successful teacher conversations and teacher corrections to its own mistakes, followed by further training to retain previous successes. That training produced a checkpoint that solved 67 of the 220 evaluation tasks.
Part 4 continued training from that checkpoint with a different kind of feedback. Instead of writing corrections for the student to copy, Qwen3.5-35B-A3B supplied probabilities for possible next tokens along the student’s sampled actions. The best resulting checkpoint solved 69 tasks. Both saved models are starting points for this post:
| Blog name | Place in the training sequence | Starting score |
|---|---|---|
| Checkpoint A | The trained model that Part 4 started from | 67/220 |
| Checkpoint B | A after probability distillation with the 35B teacher | 69/220 |
Two questions for self-distillation
The first question is whether the larger-teacher stage can be replaced. Starting again from Checkpoint A makes that comparison possible: one training branch uses the 35B teacher from Part 4, while the other uses a frozen copy of A as its teacher. Both branches begin with the same student.
The second question is whether self-distillation can improve the best result already available. That requires starting from Checkpoint B and using a frozen copy of B as its teacher. This tests an additional training stage after the larger-teacher stage.
These are separate experiments. The self-distilled version of A does not become B, even if the two happen to receive the same score. A and B always name the saved starting checkpoints in the table above.
How on-policy self-distillation works
In on-policy self-distillation, the current student generates the responses used for training. The teacher then supplies probability feedback at the positions in those responses. These probabilities describe possible next tokens, the pieces of text the model generates. Here, the frozen self-teacher and student begin with the same checkpoint. The student’s adapter changes during training; the teacher’s adapter stays fixed.
The input is a saved SQL-agent state: the instructions, question, and conversation before the next action. It contains no target action. During each training step, the student samples a new action from that state. The frozen teacher scores the possible next tokens along the sampled action, and the loss compares the student and teacher probabilities.
The following pseudocode illustrates that sequence. The scoring calls return next-token probabilities along the sampled response; they do not ask the teacher to write a replacement action.
state = saved_sql_agent_state
sampled_action = student.generate(state)
student_scores = student.next_token_scores(state, sampled_action)
teacher_scores = frozen_teacher.next_token_scores(state, sampled_action)
loss = distillation_loss(student_scores, teacher_scores)
update_student_adapter(loss)
The frozen-teacher runs used TRL’s experimental generalized knowledge distillation trainer, GKDTrainer, with lmbda=1.0. This setting generates completions from the current student during training. The prompt states were collected beforehand and stayed fixed. “On-policy” here refers to the newly sampled actions, not to recollecting complete agent conversations after every update. With lmbda=0.0, training would instead use stored completions, a different method.
This differs from correction SFT. A correction teacher can supply a replacement action for the student to imitate. Probability distillation changes the probabilities assigned along actions the student generates. It supplies no new teacher-written path through the task, although changes in those probabilities can affect later actions.
Giving the teacher additional context
Alongside the frozen-teacher comparison, a separate variant tested whether giving the small teacher additional information would help. It used TRL’s experimental SDPOTrainer. The student received the normal SQL-agent state, while the teacher side received extra training-only information: either a reference solution or feedback from the training attempt. The teacher scored the same completion the student had sampled under its normal prompt. The extra information was not used as the student’s target text and was absent during evaluation.
These privileged-context runs used teacher_model_kind="live" and sampled-token distillation, rather than the frozen-adapter GKD configuration above. The recorded distillation_weight=1.0 and a reward function returning zero for every completion supplied probability-distillation training without a task-reward preference signal.
Training data and settings
Both starting checkpoints ran on the fixed 879-task training split to collect prompt states. The datasets mixed states associated with difficult attempts and successful attempts, called anchors. A task could supply several states, so row counts exceed task counts.
| Starting checkpoint | Difficult states | Success anchors | Total rows |
|---|---|---|---|
| Checkpoint B | 1,956 | 898 | 2,854 |
| Checkpoint A | 1,983 | 864 | 2,847 |
Checkpoint B solved 289 of the 879 training tasks; Checkpoint A solved 280. The 2,847 rows from the latter were also used for Part 4’s larger-teacher comparison. These training-task counts are separate from the checkpoints’ scores on the 220 evaluation tasks.
The overall context budget was 8,192 tokens. The privileged-context runs filtered to shorter states to leave room for the teacher’s extra information.
All runs trained LoRA adapters on unsloth/Qwen3.5-0.8B. LoRA updates a small set of additional weights while keeping the base weights fixed. Rank and alpha were both 32; the batch size was one, with gradients accumulated over eight examples. Each reported training step is one optimizer update.
Two settings varied together in the GKD comparison:
| Schedule | Learning rate | Generated-token limit |
|---|---|---|
| Standard | 0.000005 | 128 |
| Lower rate, shorter responses | 0.0000005 | 32 |
The standard schedule matched the larger-teacher comparison. Because the lower-rate schedule also shortened responses, its results do not isolate the effect of learning rate.
The loss also varied. KL divergence measures a difference between probability distributions; its direction changes how mismatches contribute to the loss. The recorded forward-KL configuration used beta=0.0 and the teacher’s top 20 token probabilities plus the remaining probability mass. Reverse KL used beta=1.0 and loss_top_k=1. A mixed configuration used beta=0.5 and loss_top_k=20. These runs changed both the KL setting and the number of retained probabilities, so they do not isolate KL direction alone.
Each candidate ran on all 220 evaluation tasks. Those scores determined candidate selection; the experiment had no independent final test set. No evaluation tasks were removed.
Comparison with a larger teacher
For the first question, self-distillation from Checkpoint A reached 69/220 after five standard forward-KL steps. The larger-teacher run reached that same score after ten steps. At equal numbers of training steps, the ordering changed: self-distillation scored higher at five steps and lower at ten.
| Training steps | Self-teacher | 35B teacher |
|---|---|---|
| Starting score | 67/220 | 67/220 |
| 5 steps | 69/220 | 67/220 |
| 10 steps | 63/220 | 69/220 |
The larger teacher itself scored 96/220. Its student and the self-distilled student each peaked at 69/220 in this comparison, but required different numbers of steps. Matching the best total therefore conceals different responses to continued training.
Can self-distillation improve Checkpoint B?
The second question had a different answer. Every tested candidate from Checkpoint B scored below its starting score of 69/220. The best was the five-step standard forward-KL run at 67/220. Extending that configuration to ten steps reduced the score to 59; fifteen steps produced 61.
The lower-rate, shorter-response forward-KL run scored 63 at ten steps. Reverse KL scored 66 under that schedule, a smaller regression than the other ten-step GKD variants. As noted in the setup, these loss configurations also retained different numbers of token probabilities.
Both five-step privileged-context variants scored 65. Extending the reference-solution variant to ten steps reduced its score to 35. These runs used a learning rate of 0.0000005 and a 32-token completion limit. The additional teacher context did not improve task performance in these runs; the cause of the larger ten-step drop remains unresolved.
The appendix lists every candidate from both starting checkpoints, including the other settings tested from A.
Newly solved tasks and regressions
The five-step run from Checkpoint A gained four tasks and lost two, producing its net gain of two. Other candidates also found new successes, but none gained more tasks than it lost. The ten-step standard run from Checkpoint B found no new successes and lost ten.
| Starting checkpoint and candidate | New successes | Lost successes | Union |
|---|---|---|---|
| B: forward KL, standard, 5 steps | 5 | 7 | 74 |
| B: forward KL, standard, 10 steps | 0 | 10 | 69 |
| B: reverse KL, short, 10 steps | 2 | 5 | 71 |
| B: SDPO solution, 5 steps | 4 | 8 | 73 |
| A: forward KL, standard, 5 steps | 4 | 2 | 71 |
| A: forward KL, standard, 10 steps | 1 | 5 | 68 |
| A: reverse KL, short, 10 steps | 4 | 5 | 71 |
The union counts tasks solved by either the parent or the candidate. It is an upper bound for a system that could always select the correct answer, not the score of a tested combined model. The extra successes show where the two models differ, without demonstrating a way to combine them.
Remaining SQL and agent errors
The best self-distilled model submitted SQL on 176 tasks and solved 69. The other 107 submissions failed the hidden tests. Producing a valid action and reaching a submission therefore remained insufficient for many tasks.
| From Checkpoint A | Solved | Submitted | Repeats | Parse failures |
|---|---|---|---|---|
| Checkpoint A | 67 | 174 | 40 | 5 |
| Forward KL, standard, 5 steps | 69 | 176 | 35 | 8 |
| Forward KL, standard, 10 steps | 63 | 172 | 40 | 6 |
The five-step candidate submitted slightly more often and had fewer repeated-action stops, while parse failures increased. Recorded SQL execution errors fell from eight to six. The operational changes were mixed: fewer loops and execution errors accompanied more parse failures.
Hardware and runtime
The self-distillation runs used one NVIDIA RTX 4080 with 16GB of memory for training and evaluation. The GKD student and frozen teacher used adapters over the same small base model, allowing both to fit on that GPU. The privileged-context SDPO runs also used the small model. Neither setup required a separate large-teacher server.
A ten-step training run took about 24 minutes, while a full 220-task evaluation took about 27 minutes. Evaluation was slightly longer than training at this budget, since each task involved a multi-turn conversation and SQLite execution. The two activities therefore contributed similar amounts of runtime per candidate. No complete serving-cost comparison was recorded.
Conclusions
Self-distillation matched the best larger-teacher result from Part 4 when both started from Checkpoint A. Five steps reached 69/220 without serving the 35B model, with four new successes and two regressions. Ten steps lost that gain.
It did not extend the earlier result: none of the tested variants improved Checkpoint B. In this experiment, self-distillation worked as an alternative to one selected larger-teacher run, but did not provide another improvement afterward. The reason for the small gain remains unexplained, and whether it repeats on new tasks remains untested.
Appendix: full candidate results
“Standard” and “short” refer to the schedules in Training data and settings. Every score is out of 220 tasks; changes are relative to the starting checkpoint.
| From Checkpoint B | Steps | Score | Change |
|---|---|---|---|
| Forward KL, standard | 5 | 67 | −2 |
| Forward KL, standard | 10 | 59 | −10 |
| Forward KL, standard | 15 | 61 | −8 |
| Forward KL, short | 10 | 63 | −6 |
| Mixed KL, short | 10 | 61 | −8 |
| Reverse KL, short | 10 | 66 | −3 |
| Reverse KL, short | 5 | 62 | −7 |
| SDPO, reference solution | 5 | 65 | −4 |
| SDPO, reference solution | 10 | 35 | −34 |
| SDPO, attempt feedback | 5 | 65 | −4 |
| From Checkpoint A | Steps | Score | Change |
|---|---|---|---|
| Forward KL, standard | 5 | 69 | +2 |
| Forward KL, standard | 10 | 63 | −4 |
| Forward KL, short | 10 | 64 | −3 |
| Reverse KL, short | 10 | 66 | −1 |
References
- Part 4: On-policy probability distillation: the larger-teacher comparison used here.
- TRL GKDTrainer documentation: background on generalized knowledge distillation. This experiment used the recorded experimental trainer configuration described above.