← ResearchSmall-Model Distillation

Small-Model Distillation, Part 2: Off-Policy Soft-Label KD for a 0.8B SQL Agent

Isaac Kargar23 min read

TL;DR

The last post took a 0.8B SQL agent from 1/220 to 46/220 by copying a big teacher one token at a time. This one asks a simpler question: what if the student also got to see part of the teacher’s uncertainty, not just its final pick? I kept everything else fixed: the harness, the 220-task eval, the action schema, the 0.8B student. I only changed the signal it learned from.

I scored the teacher’s frozen traces, saved the top 20 token probabilities at each position plus a little leftover mass for everything else, and compared plain weighted hard-token training against top-k soft-label distillation. The best run hit 55/220, up from 46/220 for the hard-token student and 49/220 for the combined hard-token run.

The headline is not the interesting part. Soft labels actually hurt in two of the eight runs. And the best teacher was not the strongest one. A weaker same-family model, Qwen at 96/220, agreed with its own traces far better than the stronger model, GPT at 115/220, so its probabilities taught the student more cleanly. The win came from mixing both.

One honest caveat up front: I never had GPT’s own token probabilities. The GPT traces here are scored by Qwen, so they are a proxy, not true GPT logit distillation.

What I Wanted To Test

I am trying to make small models better at narrow tasks. Same SQL repair agent as before.

The earlier hard-token run already showed the small model can run the loop.

But hard-token training is blunt. It does not see whether the teacher was confident, whether there were plausible alternatives, or whether the next-best token was almost as good. In SQL, that uncertainty can matter around table names, column names, predicates, joins, string values, and structured action syntax.

So the question for this post was:

Does top-k soft-label distillation improve the same 0.8B SQL agent when the traces and benchmark stay fixed?

I also wanted to test two starting points. One is the base unsloth/Qwen3.5-0.8B model. The other is the Qwen3.5-0.8B hard-token adapter trained on Qwen3.5-35B-A3B 8-bit successful trajectories.

The SQL-Agent Problem

I kept the problem setup identical to the hard-token post. Each task has a user issue, a buggy SQL query, and a SQLite database. The model can choose structured actions: inspect schema, run SQL, or submit final SQL. A task succeeds only when the final submitted SQL passes the hidden tests. For a concrete broken-query example, see the chinook task in the offline teacher-trace SFT article.

BAML is still the structured-output layer around the prompt. It pins down the decision schema and builds the model request. I do not train the student on free-form teacher prose.

{"draft":"Need schema first.","output":{"action":"inspect_schema"}}
{"draft":"Test a candidate query.","output":{"action":"run_sql_query","sql":"SELECT ..."}}
{"draft":"Submit corrected SQL.","output":{"action":"submit_sql","sql":["SELECT ..."]}}

The only thing that moved was the training signal:

Eval settingValue
Held-out eval split220 SQL repair tasks
Max turns8
Local eval sequence cap8192 total tokens
Max generated tokens per model call512, reserved inside the local sequence cap
Temperature0.0
Task timeout180 seconds
Scoringdeterministic hidden SQL tests

The 8192 sequence cap is the whole prompt-plus-generation budget: system prompt, task, prior actions, schema, query results, the structured-output instructions, and the generation reserve all share it. I list max generated tokens separately because it caps the size of a single model call: the student can read a long state, but it still has to emit one compact action per turn.

The eval table uses these terminal outcomes and behavior metrics:

Category or metricMeaning
submitted but wrongThe model reached submit_sql, but the SQL failed hidden tests
SQL-error tasksAt least one model-issued SQL probe errored during the loop; this is not mutually exclusive with the final task outcome
parse failuresThe output did not parse as a valid BAML action
repeat stopsThe harness stopped the run because the model repeated actions
max-turn stopsThe model used all 8 turns without solving the task
runtime errorsProcess-level failures or timeouts

I am spelling this out because the best students are not mainly failing on JSON. They reach submit_sql just fine. The real problem is SQL judgment and recovering when a probe goes sideways.

The Training Idea

With hard-token distillation, the student gets one target token at each position. Soft-label distillation hands it a small probability distribution instead.

For example, imagine the teacher target contains a table token like ... FROM users .... Hard-token training says: the next token is users.

Soft-label training can say:

Token optionTeacher probability
users0.72
customers0.12
accounts0.06
orders0.03
everything else0.07

That is a richer signal.

Hard token versus soft labels
Hard-token training keeps one chosen target, while sparse soft-label training keeps the teacher’s top 20 alternatives and residual tail mass.

For a large language model, storing full-vocabulary probabilities at every assistant target position is expensive. So I used a practical sparse version: top 20 teacher token probabilities at each assistant target position, plus residual tail mass for all other tokens.

The tail matters. If the saved top 20 probabilities sum to 0.996, the missing 0.004 is still part of the teacher distribution. Renormalizing the top 20 to 1.0 would make the target artificially overconfident.

This is still off-policy: the student learns from traces I collected earlier and never acts while I am scoring them.

A Small Educational Code Walkthrough

I did not create a new agent environment for soft labels: same teacher decisions from the hard-token run, same BAML prompt shape, plus probability information on the assistant target tokens.

The hard-token supervised fine-tuning (SFT) row looked like this.

row = {
    "messages": conversation_before_teacher_decision + [
        {"role": "assistant", "content": canonical_decision_json(teacher_decision)}
    ],
    "teacher_draft": teacher_decision.draft,
    "teacher_action": teacher_decision.output,
}

For soft-label distillation, the row keeps the same canonical messages and teacher metadata, but adds teacher-forced token scores for assistant target positions. Teacher forcing means I feed the already-known target text to the probability model and ask, token by token, how likely that target was under the current context. The serialized data stores log probabilities, not raw probabilities, because they are numerically safer to write and train from. The IDs and position in the compact example below are illustrative; the recorded rows contain the values produced by tokenization.

row = {
    "messages": [..., {"role": "assistant", "content": target_decision_json}],
    "distillation": {
        "probability_model": "Qwen3.5-35B-A3B 8-bit",
        "top_k": 20,
        "row_weight": 1.18,
        "token_scores": [
            {
                "position": 1661,              # full-sequence token position
                "target_token_id": 412,
                "target_logprob": -0.21,
                "top_token_ids": [412, 879, 91],
                "top_logprobs": [-0.21, -2.81, -3.51],
                "top_mass": 0.90,
                "tail_mass": 0.10,
            }
        ],
    },
}

The scoring pass has one job: take a frozen row and stamp token probabilities onto it. I never ask the Qwen3.5-35B-A3B 8-bit scorer to write a new decision. I just render the prompt without the assistant answer, render the full sequence with it, and score the target part one token at a time.

prompt = chat_template(messages[:-1], add_generation_prompt=True)
full = chat_template(messages, add_generation_prompt=False)
prompt_ids = tokenize(prompt)
full_ids = tokenize(full)
if full_ids[: len(prompt_ids)] != prompt_ids:
    raise ValueError("The chat template is not prefix-stable; refusing to misalign teacher scores.")
target_ids = full_ids[len(prompt_ids):]
prefill_teacher_cache(prompt_ids[:-1])
driver_ids = [prompt_ids[-1]] + target_ids[:-1]
for i, target_id in enumerate(target_ids):
    logits = qwen_35b(driver_ids[i]).logits
    logprobs = log_softmax(logits)
    top_ids, top_logprobs = top_k(logprobs, k=20)
    save_token_score(
        position=len(prompt_ids) + i,
        target_token_id=target_id,
        target_logprob=logprobs[target_id],
        top_token_ids=top_ids,
        top_logprobs=top_logprobs,
        tail_mass=1.0 - sum(exp(top_logprobs)),
    )

The assistant-only masking stays the same as the hard-token run. Prompt tokens are context. Assistant target tokens are supervision. This matters more in an agent than in a plain question-answer task, because the prompt contains schema observations, prior tool outputs, and SQL execution results.

labels = full_ids.copy()
labels[: len(prompt_ids)] = -100

The trainer stores more tensors than ordinary SFT. Each batch still has input_ids, attention_mask, and masked labels, but it also has per-position weights, top-k token ids, top-k log probabilities, a top-k mask, and tail probabilities. Padding fills inactive positions with neutral values so the loss can ignore them cleanly.

batch = {
    "input_ids": pad(input_ids, pad_id),
    "attention_mask": pad(attention_mask, 0),
    "labels": pad(labels, -100),
    "position_weights": pad(position_weights, 0.0),
    "topk_token_ids": pad(topk_token_ids, [0] * top_k),
    "topk_logprobs": pad(topk_logprobs, [0.0] * top_k),
    "topk_mask": pad(topk_mask, [False] * top_k),
    "tail_probs": pad(tail_probs, 0.0),
}

Weighted hard-token fine-tuning is the control condition. It still trains on one chosen target token, but multiplies assistant-token cross-entropy by a confidence weight derived from the scorer’s target log probability.

z = (mean_target_logprob - dataset_mean_logprob) / dataset_logprob_std
row_weight = clip(2 / (1 + exp(-z)), 0.25, 1.75)
loss = row_weight * hard_ce

The clip range [0.25, 1.75] keeps one messy row from hijacking a batch, while still letting confident rows count a good deal more than shaky ones.

Top-k soft-label distillation adds a probability-matching term. The saved top-k probabilities are renormalized to the recorded non-tail mass, while the residual tail is treated as one additional category. After padding and shifting the batch tensors, the loss is:

import torch
import torch.nn.functional as F

# student_logits has shape [batch, sequence, vocabulary]. The logit at
# position i predicts token i + 1, so shift every target-side tensor together.
student_logits = model(input_ids, attention_mask=attention_mask).logits
shift_logits = student_logits[:, :-1, :].contiguous()
shift_labels = labels[:, 1:].contiguous()
shift_weights = position_weights[:, 1:].to(shift_logits.dtype)
shift_topk_ids = topk_token_ids[:, 1:, :].contiguous()
shift_topk_logprobs = topk_logprobs[:, 1:, :].to(shift_logits.dtype)
shift_topk_mask = topk_mask[:, 1:, :].bool()
tail_mass = tail_probs[:, 1:].to(shift_logits.dtype).clamp(0.0, 1.0)

hard_ce = F.cross_entropy(
    shift_logits.transpose(1, 2),
    shift_labels,
    reduction="none",
    ignore_index=-100,
)
student_logprobs = F.log_softmax(shift_logits, dim=-1)
student_top_logprobs = torch.gather(
    student_logprobs, dim=-1, index=shift_topk_ids
)

eps = 1e-8
valid = shift_topk_mask
teacher_top_raw_probs = torch.exp(shift_topk_logprobs).masked_fill(~valid, 0.0)
teacher_top_mass = (1.0 - tail_mass).clamp_min(0.0)
teacher_top_sum = teacher_top_raw_probs.sum(dim=-1).clamp_min(eps)
teacher_top_probs = teacher_top_raw_probs * (teacher_top_mass / teacher_top_sum).unsqueeze(-1)

student_top_probs = torch.exp(student_top_logprobs) * valid
student_tail_prob = (1.0 - student_top_probs.sum(dim=-1)).clamp_min(eps)

top_terms = torch.where(
    teacher_top_probs > 0,
    teacher_top_probs * (teacher_top_probs.clamp_min(eps).log() - student_top_logprobs),
    torch.zeros_like(teacher_top_probs),
)
topk_kl = top_terms.sum(dim=-1)
tail_kl = torch.where(
    tail_mass > 0,
    tail_mass * (tail_mass.clamp_min(eps).log() - student_tail_prob.log()),
    torch.zeros_like(tail_mass),
)
per_position_loss = hard_ce + topk_kl + tail_kl
loss = (per_position_loss * shift_weights).sum() / shift_weights.sum().clamp_min(1.0)

The student is penalized when it puts too little probability on the teacher mass. The explicit shift keeps hard_ce, topk_*, tail_mass, and shift_weights aligned with the next-token objective; shift_weights is zero for padding, context, and other inactive positions. A zero tail contributes zero instead of evaluating log(0), and the clamps keep sparse or rounded probabilities finite. In the real trainer, both cross-entropy and distillation are computed on shifted assistant target positions only.

The actual student training is still LoRA fine-tuning, which gave me a cheap way to compare many runs without fully fine-tuning the 0.8B model every time.

student = load_student("unsloth/Qwen3.5-0.8B", max_seq_length=4096)
student = add_lora_adapters(
    student,
    rank=32,
    alpha=32,
    target_modules=[
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj",
    ],
)
# training_args is a transformers.TrainingArguments instance.
# soft_label_args carries the custom sparse-KD settings used by the trainer subclass.
trainer_class = make_trainer_class(Trainer, soft_label_args)
trainer = trainer_class(
    model=student,
    args=training_args,
    train_dataset=scored_train_examples,
    eval_dataset=scored_validation_examples,
    data_collator=distillation_collator,
)
trainer.train()
adapter = save_lora_adapter(student)

The trainer was the SparseTopKSoftLabelTrainer Hugging Face Trainer subclass used by the recorded experiment. soft_label_args carries the mode, loss weights, and temperature; training_args carries ordinary TrainingArguments such as batch size, gradient accumulation, learning rate, and epochs. That is different from TRL’s online DistillationTrainer path, where the teacher can score inside the trainer. I kept the hard cross-entropy weight, the distillation loss weight, and temperature all at 1.0 on purpose: the sparse artifacts were produced at temperature 1.0 and I did not want this post to become a temperature-scaling experiment.

The final evaluation is a real rollout, not teacher forcing. The trained adapter goes back into the same SQL-agent harness and has to solve the 220 held-out tasks by acting turn by turn.

for task in held_out_eval_tasks:
    state = start_sql_agent_task(task)
    for turn in range(8):
        messages = render_baml_messages(state)
        decision = model_with_adapter.generate_action(
            messages,
            max_new_tokens=512,
            temperature=0.0,
        )
        state = execute_action_and_append_observation(state, decision)
        if state.solved or state.stopped:
            break
    results.append(score_final_submission(state))

That last step is where a lot of offline distillation ideas quietly break. Teacher forcing measures whether the student can predict the teacher’s next action in a teacher-style context. Rollout eval measures whether the student can survive its own previous actions, pick useful SQL probes, avoid repeating itself, and submit correct SQL. This post is about the second number.

Dataset And Filtering

The traces were frozen from the hard-token run, so the comparison stays focused on the supervision signal and not on a fresh data-generation run.

In that earlier data-generation pass, both teachers ran on the same 879 train tasks. GPT 5.5 medium solved 446/879 and produced 1046 SFT rows. Qwen3.5-35B-A3B 8-bit solved 394/879 and produced 1232 rows because its successful traces were longer. On the 220-task eval, GPT 5.5 medium was stronger at 115/220 versus 96/220 for Qwen3.5-35B-A3B 8-bit, so the Qwen path here is a same-family test, not the strongest-teacher path.

There are two trace sources:

Trace sourceProbability sourceWhat it means
GPT 5.5 medium successful tracesQwen3.5-35B-A3B 8-bitProxy scoring. These are not GPT 5.5 medium’s internal probabilities. They are Qwen3.5-35B-A3B 8-bit probabilities while teacher-forcing GPT’s chosen actions.
Qwen3.5-35B-A3B 8-bit successful tracesQwen3.5-35B-A3B 8-bitSame-family self-scoring. The same model family produced and scored the traces.

Given that proxy caveat, the GPT trace experiment asks a narrower question: can a same-family Qwen scorer add useful signal to GPT’s chosen actions?

Soft-label distillation flow
Frozen teacher traces are scored into sparse targets, used for LoRA training, and evaluated through the same SQL harness.

For each row, I rendered the harness’s BAML prompt, teacher-forced the known assistant target through Qwen3.5-35B-A3B 8-bit, and stored token scores only for the assistant target positions.

Before training anything, I had two questions about the scored data itself. How much of it fit the 4096-token budget? And how well did the Qwen scorer actually agree with each trace source?

Scored data alignment
The two panels compare rows fitting the 4096-token budget with the Qwen scorer’s mean target negative log likelihood.

In the right-hand panel, lower negative log likelihood is better. It means the probability model found the target actions easier to explain. That is why the Qwen self-scored data looks so different from the GPT proxy-scored data.

Scored datasetSource rowsKept at 4096Train/validationMean prompt tokensMean target tokensP95 full sequenceMean target negative log likelihoodTarget in top 20
GPT 5.5 medium traces, proxy-scored by Qwen3.5-35B-A3B 8-bit10461042990 / 5216588731710.87698.56%
Qwen3.5-35B-A3B 8-bit traces, self-scored123212111151 / 6018797934030.297100.00%

Qwen3.5-35B-A3B 8-bit found its own traces much easier to explain than the GPT traces. The mean target negative log likelihood was 0.297 for Qwen self-scored rows versus 0.876 for GPT proxy-scored rows. The target token appeared inside the stored top 20 for 100.00% of Qwen self-scored target positions and 98.56% of GPT proxy-scored positions.

The combined dataset had 2278 source rows, 2253 rows after the 4096-token filter, 2141 train rows, and 112 validation rows. Each row keeps its own scorer: the mix is data mixing across teachers, not one merged probability distribution. I shuffled the combined rows with a fixed seed before splitting so the validation set would not depend on source order.

Training Setup

The main grid was 2 x 2 x 2.

AxisValues
Student modelunsloth/Qwen3.5-0.8B
Student startbase model; Qwen3.5-0.8B hard-token trajectory adapter trained on Qwen3.5-35B-A3B 8-bit rows
Trace sourceGPT 5.5 medium successful traces; Qwen3.5-35B-A3B 8-bit successful traces
Probability modelQwen3.5-35B-A3B 8-bit
Training methodsweighted hard-token fine-tuning; top-k soft-label distillation

After those eight runs, I added two combined-teacher runs from the base model. I did not add combined warm-start runs in this post because the first eight runs already showed that warm-start behavior depended strongly on the trace source, and I wanted the combined-data test to answer a simpler question: if I start from the base model, does mixing teacher sources help?

The training recipe stayed close to the hard-token trajectory baseline: the LoRA settings from the walkthrough above, bf16 on an NVIDIA GPU, Unsloth, and FlashAttention 2 where available. I used a 5% validation split, no packing, seed 42, and a 4096-token full training sequence cap.

The residual tail mass in the loss is just a numerical correction for sparse serialized probabilities.

In the table below, Base means the unfine-tuned unsloth/Qwen3.5-0.8B model; the data labels are the trace sources defined in the Dataset section.

Student startDataMethodTrain rowsOptimizer updatesFinal validation lossTrain time
BaseGPT proxyweighted hard990372not capturednot captured
BaseGPT proxysoft label9903720.679524.5 min
BaseQwen selfweighted hard11514320.254427.6 min
BaseQwen selfsoft label11514320.490229.4 min
Hard-token trajectory warm startGPT proxyweighted hard9903720.354723.0 min
Hard-token trajectory warm startGPT proxysoft label9903720.659224.2 min
Hard-token trajectory warm startQwen selfweighted hard11514320.295627.5 min
Hard-token trajectory warm startQwen selfsoft label11514320.495029.6 min
BaseCombinedweighted hard21418040.347851.8 min
BaseCombinedsoft label21418040.568455.2 min

A small warning on this table: validation losses are useful inside one objective, but they are not directly comparable between weighted-hard and soft-label runs because the losses include different terms. The base Qwen self-scored weighted-hard run had the lowest validation loss in the table, but it solved only 40/220. The best rollout score came from the combined soft-label run, even though its validation loss was higher.

Results As Research Questions

The first result chart I made tried to put everything in one place, but it mixed two different questions, so I split the results by the actual research questions.

For scale, the teacher/baseline picture from the fixed 220-task eval was this:

ModelRole in this postSuccessSubmittedAvg turns
GPT 5.5 mediumStrong trace teacher from the hard-token data-generation pass115/2202202.53
GPT 5.4 miniHosted baseline on the same fixed eval105/2202202.14
Qwen3.5-35B-A3B 8-bitSame-family trace teacher and probability model96/2202023.57
Qwen3.5-0.8B hard-token trajectory studentWarm-start checkpoint46/2201553.48
Base Qwen3.5-0.8BNo fine-tuning1/22012.39

First: if I start from the base unsloth/Qwen3.5-0.8B model, what helps?

Base Qwen3.5-0.8B solve counts for weighted hard-token and top-k soft-label training
Held-out solved-task counts compare weighted hard-token and top-k soft-label training from the base 0.8B student.

Every training run beat the no-fine-tune baseline of 1/220. Single-source GPT 5.5 medium traces and single-source Qwen3.5-35B-A3B 8-bit traces were close, 40 to 43 solved tasks depending on method. The bigger jump came from combining teacher sources. Combined weighted hard-token fine-tuning reached 49/220, and combined top-k soft-label distillation reached 55/220. That is about a 25% solve rate, modest in absolute terms and well under the strongest teacher’s 115/220. But this series is about how far a 0.8B model can climb from a near-zero base, not about catching the teacher.

Second: if I already have the hard-token trajectory student, should I continue training it?

Warm-start Qwen3.5-0.8B solve counts after weighted hard-token and top-k soft-label continuation
Held-out solved-task counts compare weighted-hard and soft-label continuation from the 46/220 hard-token student.

Here the answer is more subtle. The starting checkpoint was already 46/220. Continuing on GPT 5.5 medium traces with weighted hard-token fine-tuning gave a small gain to 48/220, while soft-label distillation on the same GPT traces dropped to 44/220. On Qwen3.5-35B-A3B 8-bit self-scored traces, the direction flipped: weighted hard-token fine-tuning dropped to 38/220, while soft-label distillation improved to 49/220.

The part I did not expect: soft labels only paid off where the scorer agreed with the trace. The weaker teacher’s probabilities lined up with its own traces, so they helped there and actually hurt on the proxy-scored GPT traces. Alignment beat raw teacher strength.

Third: what changed in the agent loop?

Outcome decomposition for the base, hard-token, and combined soft-label SQL agents
The bars break each 220-task evaluation into solved tasks, wrong submissions, repeat stops, parse stops, and max-turn or runtime stops.

The base model almost never submitted at all. Hard-token trajectory SFT taught it to actually enter the loop and submit, but a lot of those submissions were wrong. The combined soft-label run pushed success up again, and it is the best run in the post. Even so, it submitted 171 tasks and solved 55, so 116 submitted queries failed the hidden tests. It still had 39 repeated-action stops too.

Turn count was the cleanest behavior signal I had. The combined soft-label run averaged 3.39 turns, almost the same as the hard-token warm start at 3.48, so the gain did not come from taking more steps. It came from a slightly better mix of actions: more solved submissions than the hard-token warm start, fewer wrong submissions than the combined weighted-hard run, fewer SQL errors, and fewer max-turn stops. I did not log exact generated-token counts per eval call, so I am not reporting output-token numbers here. The token stats I can report cleanly are on the training side: prompt tokens, target tokens, full-sequence length, and top-k coverage.

Exact Numbers Behind The Charts

The tables below give the values behind the baseline and trained-student charts. Submitted includes successful submissions and wrong submissions. Wrong submits is submitted - success.

Student startData sourceMethodSuccessSubmittedWrong submitsAvg turns
Base Qwen3.5-0.8BNo fine-tuningbaseline1/220102.39
Hard-token trajectory warm startQwen3.5-35B-A3B 8-bit hard-token rowsstarting checkpoint46/2201551093.48
Base Qwen3.5-0.8BGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bitweighted hard-token fine-tuning41/2201981572.60
Base Qwen3.5-0.8BGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bittop-k soft-label distillation42/2201851432.76
Base Qwen3.5-0.8BQwen3.5-35B-A3B 8-bit traces self-scoredweighted hard-token fine-tuning40/2201551153.46
Base Qwen3.5-0.8BQwen3.5-35B-A3B 8-bit traces self-scoredtop-k soft-label distillation43/2201581153.29
Hard-token trajectory warm startGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bitweighted hard-token fine-tuning48/2202021542.68
Hard-token trajectory warm startGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bittop-k soft-label distillation44/2201921483.05
Hard-token trajectory warm startQwen3.5-35B-A3B 8-bit traces self-scoredweighted hard-token fine-tuning38/2201481103.78
Hard-token trajectory warm startQwen3.5-35B-A3B 8-bit traces self-scoredtop-k soft-label distillation49/2201511023.46
Base Qwen3.5-0.8BCombined GPT proxy + Qwen self-scored rowsweighted hard-token fine-tuning49/2201751263.36
Base Qwen3.5-0.8BCombined GPT proxy + Qwen self-scored rowstop-k soft-label distillation55/2201711163.39
Student startData sourceMethodSQL-error tasksParse failuresRepeat stopsMax-turn stopsRuntime errors
Base Qwen3.5-0.8BNo fine-tuningbaseline01120800
Hard-token trajectory warm startQwen3.5-35B-A3B 8-bit hard-token rowsstarting checkpoint575440
Base Qwen3.5-0.8BGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bitweighted hard-token fine-tuning3981310
Base Qwen3.5-0.8BGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bittop-k soft-label distillation2343010
Base Qwen3.5-0.8BQwen3.5-35B-A3B 8-bit traces self-scoredweighted hard-token fine-tuning1095420
Base Qwen3.5-0.8BQwen3.5-35B-A3B 8-bit traces self-scoredtop-k soft-label distillation1045440
Hard-token trajectory warm startGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bitweighted hard-token fine-tuning3551300
Hard-token trajectory warm startGPT 5.5 traces proxy-scored by Qwen3.5-35B-A3B 8-bittop-k soft-label distillation2332410
Hard-token trajectory warm startQwen3.5-35B-A3B 8-bit traces self-scoredweighted hard-token fine-tuning1476140
Hard-token trajectory warm startQwen3.5-35B-A3B 8-bit traces self-scoredtop-k soft-label distillation1385470
Base Qwen3.5-0.8BCombined GPT proxy + Qwen self-scored rowsweighted hard-token fine-tuning1473161
Base Qwen3.5-0.8BCombined GPT proxy + Qwen self-scored rowstop-k soft-label distillation1273930

Failure Analysis

If I only looked at the combined base-model run, the story would be too clean: mixed teacher data plus soft labels wins, full stop. The split charts show the real, messier shape. The headline lesson is the counterintuitive one. A weaker teacher taught the student better than the stronger one, because its probabilities agreed with its own traces. Soft labels helped most on same-family Qwen paths and on the combined mix, and actually hurt the GPT warm-start cell.

Rollout still dominates: predicting the teacher’s next action in a teacher-style context and recovering inside your own loop are related, but they are not the same.

Validation loss shows the same gap: the combined soft-label run reached 55/220 with a higher validation loss (0.5684 vs 0.3478), and the lowest-loss run solved only 40/220. Lower teacher-forced loss was just not the same as a better agent.

The combined-teacher result is the most useful result in the post. GPT 5.5 medium traces pushed the student toward submitting more often. Qwen3.5-35B-A3B 8-bit traces gave same-family, easier-to-score trajectories, but also more loopiness. The mixed soft-label run did not inherit only the best behavior from both. It still had repeat stops. But it reached the best success score in the post.

The weighted-hard control also taught me something. Weighting a row by scorer confidence is cheaper than training against sparse probability distributions, and on GPT proxy-scored warm-start data it was the best cell. But it is still a one-token target, so top-k soft labels are the better fit when the data mix is broad enough to avoid over-specializing to one teacher style.

Hardware And Infrastructure Lessons

This post needed two different kinds of machine. Scoring meant holding Qwen3.5-35B-A3B 8-bit in memory and running it over frozen targets, so I used a high-RAM Mac where 128 GB of unified memory can keep the quantized teacher resident. It was not fast, but it let me do all the prep work locally before paying for GPU time.

Training wanted the opposite. LoRA fine-tuning is matrix heavy, and the NVIDIA GPU was much faster even with less total memory, because those tensors stayed in VRAM and ran on CUDA kernels.

WorkHardware shape that helpedObserved time
Score GPT 5.5 medium traces with Qwen3.5-35B-A3B 8-bitHigh-RAM preparation machineabout 60 min
Score Qwen3.5-35B-A3B 8-bit traces with itselfHigh-RAM preparation machineabout 41 min
Single-source student LoRA trainingNVIDIA GPUabout 23-30 min
Combined weighted-hard student LoRA trainingNVIDIA GPUabout 52 min
Combined soft-label student LoRA trainingNVIDIA GPUabout 55 min

The lesson I keep relearning: capacity and throughput are different problems. My workflow settled into doing everything I could on the Mac, and only firing up the remote GPU once the data was ready.

What I Learned

Soft-label distillation is worth testing for agent fine-tuning, but it needs rollout evaluation: token-level supervision is still local, and it does not automatically teach recovery, loop avoidance, or better SQL probing.

Proxy soft labels should be labeled honestly: the GPT rows here are Qwen-scored, not true GPT logit distillation.

A same-family teacher can be the better teacher even when it is not the strongest one. Alignment with the trace mattered more than the eval score.

Sparse top-k plus tail is a practical compromise: top 20 plus residual tail kept enough signal to make the target token visible almost everywhere.

Mixing teachers helped more than choosing one winner.

Conclusion

The answer is yes, but with conditions. Top-k soft-label distillation improved the 0.8B SQL agent in the best setup: combined GPT 5.5 medium proxy-scored rows plus Qwen3.5-35B-A3B 8-bit self-scored rows, trained into the base unsloth/Qwen3.5-0.8B student, reached 55/220. That beats the off-policy hard-token trajectory student at 46/220 and the combined hard-token run at 49/220.

But this is not a simple “soft labels always win” recipe. Teacher source, student initialization, probability alignment, and rollout behavior all mattered.

That is why I like this benchmark for the series. The student can look good under teacher forcing and still fail when it has to act. The next useful experiment should keep the same harness and ask whether a new training idea improves the actual loop, not only the offline loss.

Your company. Your workspace.

Bring Nablo to your company.

Talk with us about the data your team works with and what you want Nablo to do.