← ResearchSmall-Model Distillation

Small-Model Distillation, Part 4: On-Policy Probability Distillation for a 0.8B SQL Agent

Isaac Kargar21 min read

TL;DR

I started this post from the best hard-token correction-family SQL agent I had so far: unsloth/Qwen3.5-0.8B, evaluated at 67/220 on the fixed eval set. That parent model came from supervised fine-tuning, or SFT, on hard assistant/action tokens, first from teacher traces, then from DAgger-style expert corrections, and finally from retention-heavy anti-forgetting consolidation. Here I wanted to test a different signal: keep the same 0.8B student, keep the same SQL harness, but use a larger same-family Qwen/Qwen3.5-35B-A3B teacher to score the student’s own sampled action tokens with log probabilities.

The clean result was small but real. The best on-policy probability distillation run reached 69/220, a +2 task gain over the 67/220 parent. It used TRL’s standard DistillationTrainer, lmbda=1.0, a forward-Kullback-Leibler-style setting, top-20 teacher probabilities, 10 optimizer steps, and full 220-task evaluation. The negative result is just as important: longer runs, a lower learning rate, a second fresh-state round, and a different random seed all got worse. Probability distillation helped as a careful nudge, not as a new capability source.

One honest caveat up front: the teacher here scores tokens the student already samples. If the student never samples a good SQL strategy from a given state, token-level teacher probabilities cannot invent that missing behavior. The gain came from cleaning up the student’s local action distribution, not from teaching it entirely new SQL.

What I Wanted To Test

The question was simple: can a larger same-family teacher improve a small SQL tool-use agent by scoring the tokens the student already produces? This is different from asking the teacher to fix the task. In the earlier DAgger-style expert-correction SFT article, the teacher wrote a target action and the student learned to imitate it. In this post, the teacher does not write the action. The current student samples an assistant/action span, and the larger teacher scores those exact tokens.

That distinction matters because the learning signal is softer. A hard target says, “write this next action.” A probability target says, “given this state and this sampled action, here is how the larger model distributes probability over the action tokens.” For a tiny SQL agent, I wanted to see whether this dense token-level feedback could clean up behavior without changing the benchmark, parser, scorer, harness, action schema, train split, eval split, max turns, or eval budget.

The SQL-Agent Problem

The task is still the same SQL-agent benchmark from the earlier posts. The fixed eval set has 220 tasks, and a task counts as solved only when the submitted SQL passes the scorer.

BAML is still the structured-output layer around the harness, untuned in this post. A typical action still looks like this:

{"draft":"Inspect the schema before writing SQL.","output":{"action":"inspect_schema"}}
{"draft":"Test the join and filter.","output":{"action":"run_sql_query","sql":"SELECT ..."}}
{"draft":"Submit the corrected query.","output":{"action":"submit_sql","sql":["SELECT ...;"]}}

The harness records terminal outcomes such as parse failure, repeated-action stop, max-turn stop, runtime error, and submitted-but-failed. A SQL execution error is a diagnostic event that can occur before the episode reaches one of those terminal outcomes. These measures matter because an agent can become cleaner without becoming more correct. A model can submit more often, repeat itself less, and still submit the wrong query.

The Training Idea: Off-Policy Versus On-Policy

In the offline teacher-trace SFT article and the offline soft-label KD article, both the prompt state and the assistant target came from the teacher. The teacher ran the harness, left successful traces, and each teacher action became a supervised row. The soft-label article added a richer target, top-k teacher probabilities at each position, but the rows were still frozen teacher traces. The student never acted while those targets were being scored. That is off-policy distillation: the data comes from another policy, and it does not change between training steps.

In the DAgger-style expert-correction SFT article, the prompt state moved on-policy. The student rolled out, reached its own states, and the teacher wrote a continuation from there. But the target was still a hard token: the teacher’s chosen action text. So the state distribution was the student’s, while the target was the teacher’s.

This post moves the target on-policy too. During training, the current student samples the assistant/action span. This is on-policy probability distillation.

Off-policy (teacher-trace SFT): teacher state   -> teacher hard target
Off-policy (soft-label KD):    teacher state   -> teacher soft-label target
  On-policy state (expert correction): student state  -> teacher hard correction
  On-policy target (this post): student state  -> student-sampled tokens, teacher-scored

The prompt/state pool is collected from the SQL harness before training. Those rows are static during a training run. The on-policy part is the completion span sampled by the current student inside the trainer, which changes as the LoRA adapter updates. I also tested one refreshed-state second round, where I collected a new prompt/state pool from the 69/220 OPD model and trained again from that model.

This is the conceptual loss shape:

prompt = sql_agent_state_before_next_action
sampled_tokens = current_student.sample(prompt)
teacher_logprobs = qwen_35b_teacher.score(prompt, sampled_tokens)
student_logprobs = current_student.score(prompt, sampled_tokens)
loss = distillation_loss(student_logprobs, teacher_logprobs)

The important phrase is “same sampled tokens.” The student generates a completion, and the teacher must score that exact completion under the same prefix. If the teacher writes a replacement action, the method becomes hard-token correction SFT, which belongs to the expert-correction article. If the student trains on a frozen rollout buffer for many epochs without refreshing or sampling from the current policy, the method becomes more like the offline soft-label KD article. Here the teacher is a scorer, not a repair model.

Forward-KL Versus Reverse-KL, And The Three Trainer Knobs

The TRL DistillationTrainer exposes three knobs that decide the shape of the on-policy update.

lmbda controls whether the completion tokens come from the current student or from a stored buffer. lmbda=1.0 means the trainer samples completions from the current student policy at every step. That is what makes this on-policy. A lower lmbda would mix in stored completions, which moves the method back toward off-policy. Every reportable run in this post used lmbda=1.0.

beta controls the direction of the Kullback-Leibler, or KL, divergence in the distillation loss. Forward KL, which corresponds to beta=0.0 in this trainer, measures the expected divergence from the teacher distribution to the student distribution. It penalizes the student whenever the teacher has probability mass that the student does not cover, so the student is pushed to spread mass across all of the teacher’s alternatives. That is why forward-KL runs can use top-20 teacher probabilities with tail mass: the student is being asked to respect the full shape of the teacher distribution, not just its peak.

Reverse KL, which corresponds to beta=1.0 here, measures the divergence from the student distribution to the teacher distribution. It penalizes student mass in regions where the teacher assigns little probability. In these runs, top-1 teacher scoring made the sparse target concentrate on the teacher’s highest-probability token, so the update behaved like a soft version of “copy the teacher’s pick.” That concentration comes from this top-1 configuration; it is not a universal description of reverse KL.

The results section shows that the forward-KL top-20 setting was the only one that beat the parent.

loss_top_k controls how many teacher alternatives survive in the sparse probability target. Full-vocabulary logits would be the most complete signal, but they are expensive to move across the network and expensive to train against. With an external teacher served over HTTP, the trainer asks the teacher for the top-k token logprobs at each sampled position. loss_top_k=20 keeps the top 20 plus a residual tail mass for everything else, which is the same sparse-plus-tail compromise used in the offline soft-label KD article. loss_top_k=1 keeps only the teacher’s single most likely token, which pairs naturally with the reverse-KL setting.

Educational Code Walkthrough

The training row is a prompt/state row, not a supervised example with an assistant target. The row ends right before the next assistant action and contains no stored completion. A simplified row looks like this:

row = {
    "id": "TRAIN_123_turn_2_difficult_state",
    "messages": [
        {"role": "system", "content": sql_agent_instructions},
        {"role": "user", "content": task_and_current_observation},
        {"role": "assistant", "content": previous_action_if_any},
        {"role": "user", "content": previous_tool_observation_if_any},
    ],
    "metadata": {
        "task_id": "TRAIN_123",
        "db_id": "books",
        "task_category": "bugfix",
        "state_role": "difficult_state",
        "turn": 2,
        "source": "verified_student_state_pool",
        "row_contract": "prompt_state_only_no_assistant_target",
        "prompt_tokens": 1881,
    },
}

There is no teacher_action, no target, no token_scores. This is the hard boundary that separates this post from the expert-correction article: those rows carried a verified teacher-written target, while these rows carry only the state.

The mask keeps the objective honest: the loss belongs only on the sampled assistant/action tokens, never the system, task, or observation context:

system/user/tool-observation tokens -> context only -> loss mask = 0
current sampled assistant/action tokens -> trainable span -> loss mask = 1
teacher logprobs -> scores for the same sampled assistant/action tokens

During training, the trainer does three things per row. First, it renders the prompt and samples a completion from the current student. Second, it sends the prompt plus sampled token IDs to the external teacher server and receives teacher logprobs over the sampled positions. Third, it computes the distillation loss on the sampled positions only.

# Simplified on-policy sampling and scoring inside the trainer.
import torch

prompt_ids = tokenize(render_chat_template(row.messages, add_generation_prompt=True)).unsqueeze(0)
generated_ids = current_student.generate(
    prompt_ids,
    max_new_tokens=128,
    temperature=1.0,
    top_p=0.95,
)
# Hugging Face generate returns the prompt followed by newly sampled tokens.
assert prompt_ids.ndim == generated_ids.ndim == 2
sampled_ids = generated_ids[:, prompt_ids.shape[-1] :]
if sampled_ids.shape[-1] == 0:
    raise ValueError("The student returned an empty completion.")
full_ids = generated_ids

# Ask the external teacher to score the exact sampled tokens.
teacher_response = teacher_server.post("/get_sequence_logprobs/", json={
    "sequences": [full_ids[0].tolist()],
    "prompt_lengths": [prompt_ids.shape[-1]],
    "top_logprobs": 20,
    "temperature": 1.0,
})
teacher_top_ids = teacher_response["top_token_ids"]
teacher_top_logprobs = teacher_response["top_logprobs"]

# Student scores the prediction positions that produce the sampled tokens.
student_logits = current_student.forward(full_ids).logits
student_logprobs = torch.log_softmax(student_logits, dim=-1)
prompt_len = prompt_ids.shape[-1]
sampled_len = sampled_ids.shape[-1]
if prompt_len == 0:
    raise ValueError("A sampled completion needs at least one prefix token.")
prediction_positions = torch.arange(
    prompt_len - 1,
    prompt_len + sampled_len - 1,
    device=full_ids.device,
)
student_sampled_logprobs = student_logprobs[:, prediction_positions, :]
assert student_sampled_logprobs.shape[1] == sampled_len

# Loss only on sampled positions, not on the prompt.
loss = distillation_loss(
    student_sampled_logprobs,
    teacher_top_ids,
    teacher_top_logprobs,
    beta=0.0,
    loss_top_k=20,
)

The real TRL DistillationTrainer handles batching, padding, KV-cache reuse, gradient accumulation, and the LoRA backward pass. The position range above starts at prompt_len - 1 because the logit at that position predicts the first sampled token; the final sampled token is predicted at prompt_len + sampled_len - 2.

The teacher and student are from the same model family, which keeps token scoring auditable. Both use the same Qwen vocabulary, so the token IDs the student samples have the same interpretation when the teacher scores them.

Algorithm Walkthrough

Algorithm: Larger-Teacher On-Policy Probability Distillation for a SQL Agent
Input:
  M0: Qwen3.5 0.8B SQL agent from hard-token correction SFT, evaluated at 67/220
  T: Qwen/Qwen3.5-35B-A3B probability teacher, served as mlx-community/Qwen3.5-35B-A3B-8bit
  train_tasks: fixed 879-task SQL-agent train split
  eval_tasks: unchanged 220-task SQL-agent eval split
  harness: unchanged parser, action schema, SQL tools, scorer, and stop rules
  trainer: TRL DistillationTrainer with lmbda=1.0 (current-student sampling enabled)
Output:
  M1: Qwen3.5 0.8B adapter trained with larger-teacher token probabilities
  eval_1: full 220-task eval report
Data:
  Prompt/state rows collected from M0 running in the unchanged SQL harness
  Rows end before the next assistant action and contain no assistant target
Runtime flow:
  1. Run M0 on train_tasks in the unchanged SQL harness.
  2. Save prompt/state rows before assistant turns.
  3. Mix difficult states with successful-anchor states from the same parent.
  4. Start TRL probability-distillation training from M0.
  5. For each prompt/state row, sample assistant/action tokens from the current student.
  6. Ask T to score the exact sampled tokens under the same prefix.
  7. Mask system, user, and environment-observation tokens.
  8. Update only sampled assistant/action tokens with the distillation loss.
  9. Save the adapter.
  10. Evaluate the adapter on eval_tasks with the unchanged harness.
Method boundary:
  This is probability distillation, not teacher-written correction, RL, DPO,
  adapter mixing, or a harness change.

The hardware flow is different from the earlier posts: the student trained on a CUDA GPU, while the 35B teacher ran on my Mac through MLX and served logprobs over HTTP.

The CUDA student samples action tokens while the same-family MLX teacher on a Mac returns probabilities for those exact tokens
The CUDA student samples actions while the same-family MLX teacher scores those token IDs and returns log probabilities.

Dataset And Filtering

I collected the first prompt/state pool by running the 67/220 parent over all 879 train tasks in the unchanged SQL harness. The mix was roughly 2:1 difficult states to successful anchors; the exact counts are in the table below.

The two row types do different things. Difficult states are prompt/state rows taken from trajectories where the parent eventually failed: wrong submissions, SQL execution errors, repeated-action stops, and near-max-turn states. These expose the student’s current failure distribution, which is exactly where a teacher nudge could help. Successful-anchor rows are prompt/state rows taken from trajectories the parent solved. They do not contain the parent’s actions as targets. They are just states from good trajectories, included so the on-policy update does not drift away from behavior that already works. This is the same retention intuition used in the expert-correction article, adapted to the prompt/state-only contract: the anchor rows’ presence in the pool keeps the sampled completions close to states the student already handles well.

The rows were filtered by rendered prompt length, not by target length, because there is no stored target. The maximum sequence length was 8,192 tokens. At eval time the parent used the same 8-turn limit and 512-token generation budget as the earlier posts. Most runs, including the selected run, used a 128-token assistant budget during OPD training, and Run A tested a smaller 64-token cap. The final parent-state pool had prompt-token statistics of min 580, median 1,938, p90 3,016, p95 3,276, and max 5,778. No rows were dropped for being malformed, having assistant targets, or being too long.

Pool sourceTrain tasksParent train successesTotal rowsDifficult-state rowsSuccessful-anchor rows
First pool from 67/220 parent8792802,8471,983864
Second fresh-state pool from 69/220 OPD model8792892,8541,956898

I also collected a second fresh-state pool from the best 69/220 OPD model to test whether a second on-policy round would compound. The pool was useful as an experiment, but the trained second-round model fell to 64/220.

The on-policy training pool mixes difficult student states with successful-anchor states from the same parent
Each training pool combines difficult states with successful-anchor states collected from its parent checkpoint.

Training Setup

The best run used LoRA on the attention and MLP projections; every setting is in the table below.

SettingValue
Studentunsloth/Qwen3.5-0.8B
Starting checkpointExpert-correction SFT adapter at 67/220
Probability teacherQwen/Qwen3.5-35B-A3B, served as mlx-community/Qwen3.5-35B-A3B-8bit
TrainerTRL 1.6.0 trl.experimental.distillation.DistillationTrainer
lmbda1.0 (current-student sampled completions)
beta0.0 (forward-KL-style update)
loss_top_k20 (top-20 teacher probabilities plus tail mass)
LoRA rank / alpha / dropout32 / 32 / 0.0
LoRA target modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Batch / grad accumulation1 / 8
Learning rate5e-6
Optimizer steps10 (selected run)
Max sequence length8,192
Max completion length128 (Run A used 64)
Precisionbf16
Transform / trainer stackTransformers 5.12.1, PyTorch 2.10.0, PEFT 0.19.1

The results sweep is the cleanest way to see the forward-versus-reverse distinction: same student, same teacher, same prompt/state pool, different divergence direction and different target sparsity.

Results As Research Questions

Did Larger-Teacher Probability Distillation Improve The 67/220 Parent?

Yes, but only by a small amount. The best run reached 69/220.

Solved-task counts and behavior counters for the larger-teacher on-policy distillation sweep
The sweep compares full-evaluation success counts across divergence direction, top-k setting, update budget, and seed.

This is the exact table behind the chart. I include the behavior counters because the success score alone hides how the policy moved. Submitted means the agent reached submit_sql; repeated and parse counts are terminal outcomes, while SQL errors count diagnostic tool failures that can occur before the episode ends. A submitted-but-wrong task is Submitted - Eval success, so it is separate from SQL-error events.

RunMethod settingEval successSubmittedRepeated stopParse failuresSQL errors (diagnostic)Avg turns
Starting parentHard-token correction-family SFT checkpoint67/22017440583.21
AReverse-KL-style, top-1 teacher token, 5 steps, 64-token cap65/22017142773.25
BReverse-KL-style, top-1 teacher token, 5 steps, 128-token cap66/22017439663.23
CReverse-KL-style, top-1 teacher token, lower learning rate, 5 steps62/22017043753.16
DForward-KL-style, top-20 teacher distribution, 5 steps67/22017142683.21
EForward-KL-style, top-20 teacher distribution, 10 steps69/22017438783.22
FForward-KL-style, top-20 teacher distribution, 15 steps63/22017539693.28
GForward-KL-style, top-20 teacher distribution, 12 steps65/220172417113.29
HIterative fresh-state second OPD round from the 69/220 model, 5 steps64/220176386143.23
IForward-KL-style, top-20 teacher distribution, 10 steps, seed 761/220170436103.30

The selected model is Run E. Because the same 220-task evaluation set was used to compare and select the sweep, this is a selection-aware result rather than an untouched final estimate. I would not call the method broadly robust from this evidence. My read is that this kind of probability distillation can nudge a good agent, but the update budget is fragile. It is easy to move the policy away from useful behavior.

The KL-direction pattern is the cleanest signal in the sweep. All three reverse-KL-style runs (A, B, C) scored below the parent, ranging from 62 to 66. All five forward-KL-style runs that used the same 128-token cap and the standard learning rate landed between 63 and 69. The forward-KL setting gave the student a richer target, the teacher’s full top-20 distribution, and that richer signal was what produced the only gain over the parent. The reverse-KL top-1 setting collapsed the target to one token and consistently pulled the student slightly below where it started. That is consistent with the intuition: sharpening a small agent toward a single teacher token can overwrite fragile skills faster than it adds useful ones.

The step-count pattern is the other clean signal. Five steps of forward-KL tied the parent at 67. Ten steps improved to 69. Twelve steps fell to 65, and fifteen steps fell to 63. There is a narrow window where the update helps, and past that window the policy drifts. This is the same fragility seen in the expert-correction article: more updates did not keep helping.

The table reports the completed full-evaluation runs that started from the same 67/220 parent and followed this larger-teacher OPD contract.

Did The Failure Shape Improve?

The best run solved two more tasks than the parent, but the failure shape did not transform.

Failure-shape comparison between the 67-task parent and larger-teacher OPD candidates
The comparison shows submitted-but-failed tasks, repeat stops, parse failures, and SQL execution errors for selected OPD runs.

That is a small local change, not a new SQL reasoning capability. Repeated-action stops went down by two and submitted-but-failed tasks went down by two, but parse failures rose slightly from 5 to 7. SQL execution errors held flat at 8. The model became marginally cleaner in loop control without becoming more accurate in SQL judgment. This is the same pattern that recurred across the series: a small score gain can come from fewer repeated actions and slightly better submit decisions, while the dominant submitted-but-failed failure mode barely moves.

Did The Best OPD Model Solve Different Tasks?

Yes, a little. The best OPD model and the starting parent overlapped on 61 solved eval tasks. The OPD model gained 8 tasks that the parent missed, but lost 6 tasks the parent had solved. The oracle union would be 75, which is useful analysis but not a deployable single-model score.

OverlapTasks
Both solved61
Only best OPD model solved8
Only starting parent solved6
Neither solved145
Oracle union75
Tasks newly solved and lost by the best OPD candidate relative to its parent
The best OPD candidate gained eight parent misses and lost six parent wins on the same 220 tasks.

This overlap is the part that keeps showing up across the series. In the expert-correction article, the 64/220 and 67/220 checkpoints also traded tasks rather than simply accumulating them. Small training changes often move the solved-task set rather than simply adding solved tasks. For this post, I am not turning that into adapter mixing or selector logic, because that would be a different method. The lesson here is narrower: probability distillation changed the policy enough to trade a few tasks, and the net trade was +2.

Failure Analysis

If I only looked at the best run, the story would be too clean: forward-KL top-20 at 10 steps wins, full stop. The sweep shows the real, messier shape, and it matches what the other posts found.

The gain was narrow and seed-sensitive. The selected run reached 69/220, but the same 10-step forward-KL top-20 recipe with seed 7 reached 61/220. That is an 8-task swing from a seed change alone, on a 220-task eval. I would not claim the method is robust from a single winning run. What I can say is that, under the right settings, the forward-KL top-20 update produced a small real gain, and the reverse-KL top-1 setting never did.

Submitted-but-wrong SQL is still the dominant failure. The best OPD run submitted 174 times and solved 69, which means 105 submitted queries still failed the hidden tests. That is only two fewer than the parent’s 107. The probability distillation update changed how the student distributes mass over action tokens, but it did not meaningfully change whether the student’s final SQL was correct. This is the same wall every post in the series has hit: the student knows how to run the loop, but its SQL judgment is not strong enough to pass the scorer on the hard tasks.

The second fresh-state round is especially instructive. Run H started from the 69/220 model, collected a new prompt/state pool from it, and trained another 5 steps. That model scored 64/220. It kept repeated-action stops low at 38, but SQL execution errors rose to 14, the highest in the sweep. More confident action was not automatically better action. The second round pushed the policy further from the parent, and some of that push turned into faster, wronger SQL. This echoes the expert-correction article’s finding that correction-only iteration did not compound: once the parent is already strong, a second round of the same recipe can do more harm than good.

The update-budget pattern is the clearest negative result. Forward-KL top-20 rose from 67 at five steps to 69 at ten, then fell to 65 at twelve and 63 at fifteen. The policy moved past the useful update window and into territory where the LoRA adapter was overwriting fragile skills. The on-policy samples at steps 12 and 15 came from a student that had already drifted, so the teacher was scoring worse completions than it was at step 5. This is the compounding-drift risk of on-policy methods: each step’s data depends on the current policy, so a few bad steps can cascade.

My interpretation ties back to the method boundary. Hard-token correction in the expert-correction article can show the student a better action it never would have sampled. Probability distillation in this post scores only the actions the student already samples. If the student’s current policy never puts meaningful mass on the right SQL strategy in a given state, token-level teacher probabilities may not be enough to invent that missing behavior. The forward-KL top-20 setting helps because it spreads the student’s mass toward teacher-likely alternatives, but it still cannot steer the student toward a token the student essentially never considers. That is why the gain was +2 and not +10.

Hardware And Infrastructure Lessons

The Mac-plus-GPU split was the practical path, and it is different from the infrastructure pattern in the earlier posts. In the offline soft-label KD article, teacher scoring was a one-time batch pass over frozen traces before GPU training. Here the teacher is a live scorer that responds during training because the completions are sampled on-policy and are not known ahead of time.

The Mac has enough unified memory to serve the 35B 8-bit teacher with MLX. The remote NVIDIA GPU has faster CUDA kernels for LoRA training, but not enough VRAM to host both student training and the 35B teacher comfortably. Splitting them let the GPU train the 0.8B student while the Mac returned teacher log probabilities. Both sides used the same Qwen vocabulary, and the trainer checked that the teacher scored the sampled token span before updating the adapter.

The best 10-step run took about 571 seconds of training time. Each reportable experiment also needed a full 220-task harness evaluation, which is slower because each task is a multi-turn agent episode with SQLite execution. The sweep therefore stayed small enough to compare update budgets and divergence settings without treating optimizer time as the main cost.

What I Learned

The clean answer is that larger-teacher probability distillation helped, but only slightly. Starting from a strong 67/220 hard-token SFT parent, the best same-family 35B teacher scoring run reached 69/220. That is a real result, but it is not strong enough to claim that OPD solved the remaining bottleneck.

The more useful lesson is about what the method can and cannot teach: it can make the local action distribution more teacher-like, but the final task still depends on exploration, schema understanding, SQL judgment, and not overwriting fragile skills.

The forward-KL-versus-reverse-KL distinction was the clearest method-level signal. Forward-KL top-20, which asks the student to cover the teacher’s full top-20 distribution, produced the only gain. Reverse-KL top-1, which sharpens the student toward the teacher’s single pick, consistently underperformed. For a small agent that already has fragile but useful behavior, spreading mass toward alternatives was safer than collapsing onto one token.

I also learned that more OPD is not automatically better. Ten steps beat five steps, but twelve and fifteen steps got worse. A second fresh-state round also got worse. The method is sensitive enough that I would treat it as a short, carefully evaluated nudge, not a long training recipe to run blindly. The compounding-drift risk is real: each step trains on samples from the current policy, so a few bad steps can pull the policy into a worse region and stay there.

Conclusion

This post answered the narrow question I cared about: can a larger same-family logprob teacher improve the best 0.8B SQL agent by scoring the student’s own sampled tokens? Yes, but modestly.

The result is valuable because it is clean. Same student, same harness, same eval, no teacher corrections, no RL, no adapter mixing, and a standard TRL probability-distillation trainer. My takeaway is that larger-teacher OPD is useful as a precise post-training tool, but for this SQL agent it is not enough by itself. The remaining gains probably need methods that add missing successful behavior while preserving what the agent already knows. Probability distillation can clean up the action distribution the student already has, but it cannot teach the student a strategy it never samples.

That question sets up the next post in the series: instead of a larger teacher, can the 0.8B agent improve by teaching itself, scoring its own sampled tokens with on-policy probability feedback from its own checkpoints?

Your company. Your workspace.

Bring Nablo to your company.

Talk with us about the data your team works with and what you want Nablo to do.