Small-Model Distillation, Part 1: Offline Teacher-Trace SFT for a 0.8B SQL Agent
TL;DR
I wanted to see whether a very small language model could learn to act inside a real SQL tool-use loop. The model was not asked to only write a final answer. It had to inspect a SQLite database, run SQL probes, read tool observations, and submit corrected SQL that passed hidden deterministic tests.
The first method in this series is the simplest useful one: offline teacher-trace hard-token SFT. A stronger teacher runs the fixed harness first. I keep only successful train trajectories. Then each teacher decision becomes a supervised fine-tuning (SFT) row of the form conversation so far -> next structured decision.
| Role | Result |
|---|---|
| Main 0.8B student | unsloth/Qwen3.5-0.8B moved from 1/220 to 44/220 with GPT 5.5 medium rows, and to 46/220 with Qwen3.5-35B-A3B 8-bit rows. |
| Comparison students | unsloth/Qwen3.5-2B reached 58/220 at best; LiquidAI/LFM2.5-8B-A1B reached 50/220 at best. |
| Strong baselines | Qwen3.5-35B-A3B 8-bit solved 96/220, GPT 5.4 mini solved 105/220, and GPT 5.5 medium solved 115/220. |
The answer is not “the small model became as good as the teacher.” It did not. The better answer is: hard-token trajectory SFT transferred the agent protocol, but not teacher-level SQL judgment. The students learned to operate the harness and submit much more often, but most of the remaining failures moved downstream into wrong SQL submissions and loop-control mistakes.
The Method
The reusable idea in this post is a hard-token SFT contract. The SQL-agent code runs the teacher in the harness, keeps successful trajectories, and turns each assistant decision into a messages -> assistant target row. Training masks the context and learns only the teacher-written hard target with LoRA adapters. The complete teaching example below shows the label construction and next-token shift.
What I Wanted To Test
I am interested in small language models that become good at a narrow, useful task. Not a general chat model. Not a leaderboard toy. A small model that can sit inside a real workflow and do one thing well enough to matter.
For this first post, I started with a SQL repair agent. Each task gives the model a user issue, a buggy SQL query, and a SQLite database. The model can inspect the schema, run SQL queries, read observations, and eventually submit corrected SQL. The final score is deterministic: the submitted SQL either passes hidden tests or it does not.
The question was:
Can a tiny model learn the behavior of a stronger SQL tool-use agent from saved teacher trajectories?
That wording matters. I was not only asking whether the model can write SQL-looking text. I wanted the model to learn the loop: choose an action, respect the schema, read the observation, decide what to do next, and stop when it has enough evidence.
I kept the benchmark, action schema, parser, stop rules, and eval split fixed. Small-agent results are easy to distort by changing the harness, so I did not tune the environment per model.
The SQL-Agent Problem
Here is the kind of task the model sees:
Database: chinook
User issue: I want to find the latest track_id and use that id to filter records in the track table.
Buggy SQL:
WITH vars AS (SELECT COUNT(*) AS vars_id FROM track)
SELECT * FROM track WHERE track_id = vars_id
The bug is small but meaningful. COUNT(*) is not the latest id. The intended repair is closer to MAX(track_id). The buggy SQL gives useful clues about tables and columns, but it can also anchor the model to the wrong operation.
That gives the model two jobs:
- Operate the agent protocol correctly.
- Choose the right SQL repair.
A weak model can fail before doing SQL reasoning at all. It can inspect schema, inspect schema again, emit invalid JSON, or repeat an unproductive action until the harness stops it. A better model can drive the harness but still submit SQL that fails the hidden tests. Those are different failure modes, and I wanted the experiment to separate them.

The benchmark source was birdsql/six-gym-sqlite.
Concretely, the benchmark filter was category == "Query" and db_id in {"netflix", "movie_3", "books", "chinook"}. I chose this fixed slice to keep the first experiment narrow while preserving several domains, then split inside each database so train and eval kept the same domain mix.
| Setting | Value |
|---|---|
| Task category | Query |
| Databases | netflix, movie_3, books, chinook |
| Source rows scanned | 5000 |
| Candidate rows after filtering | 1099 |
| Train split | 879 tasks |
| Eval split | 220 tasks |
| Split seed | 42 |
The split by database was:
| Database | Candidate tasks | Train tasks | Eval tasks |
|---|---|---|---|
books | 282 | 226 | 56 |
chinook | 251 | 201 | 50 |
movie_3 | 273 | 218 | 55 |
netflix | 293 | 234 | 59 |
| Total | 1099 | 879 | 220 |
The model can choose only three structured actions:
{"action": "inspect_schema"}
{"action": "run_sql_query", "sql": "SELECT ..."}
{"action": "submit_sql", "sql": ["SQL statement 1", "SQL statement 2"]}
BAML, or Boundary Markup Language, is the structured-output layer around the prompt contract. It defines the action schema and renders the model request. Hosted teacher and baseline calls use the BAML structured parsing path; local HF/PEFT student eval decodes model text and the harness validates it with parse_decision. In both paths, the accepted target is canonical decision JSON, not free-form teacher prose.
The loop is simple:
- Build messages from the task.
- Ask the model for one structured action.
- Parse and validate the action.
- Execute the action in SQLite when needed.
- Append the observation.
- Repeat until a stop condition below.
The stop categories are part of the result, not just logging noise.
| Stop reason | What it means |
|---|---|
submitted | The model submitted final SQL; the SQL either passed or failed hidden tests |
parse_failure | The model did not produce a valid structured action |
repeated_action | The model repeated an action that the harness considered unproductive |
max_turns | The model kept acting but never reached a valid final submission |
runtime_error | The model call or harness call failed during the task |
For this task, formatting is not cosmetic. A human can understand “I should inspect the schema first,” but the harness needs {"action":"inspect_schema"}. I did not use keyword matching or an LLM judge to rescue malformed actions.
The Training Idea
This post is about post-training: shaping a base model’s behavior for a use case, not pretraining.

Knowledge distillation means using a stronger teacher model to transfer useful behavior into a smaller student model. For agents, the important question is not only “which teacher?” It is also:
Which part of the teacher behavior becomes supervision?
| Distillation signal | What the student learns from | Why it matters |
|---|---|---|
| Hard labels | The teacher’s chosen output tokens | Simple SFT: given this input, produce this exact target. |
| Soft labels / logits | The teacher’s probability distribution over next tokens | Preserves uncertainty over alternatives instead of only the winning token. |
| Feature distillation | Internal teacher activations or hidden states | Tries to align representations, but needs direct teacher internals. |
| Final-answer distillation | Only the completed teacher answer | Useful for answer-only tasks, weak for agents because it discards the process. |
| Trajectory distillation | Intermediate actions, observations, and final answer | Teaches the policy: inspect, query, react, and submit. |
| Reward or RL-style methods | A scalar outcome after rollout | Optimizes behavior through environment feedback instead of direct imitation. |

This first post uses offline hard-token trajectory distillation. Offline means the teacher runs first and training happens later from saved traces. Hard-token means the student trains on the exact tokens the teacher chose, not the teacher’s probability distribution. Trajectory means the training rows come from intermediate tool-use decisions, not only final SQL.
So the unit of learning is:
Full conversation state before the teacher decision -> canonical next executable decision.
This is still a local imitation objective. It does not directly optimize task success under the student’s own rollouts. That limitation becomes important in the results.
A Small Educational Code Walkthrough
The runnable example below is a small PyTorch version of the training contract. It uses whitespace tokens and a tiny causal model so it can run without downloading a language model. The actual experiment used Qwen’s chat template, canonical BAML decisions, and LoRA adapters; this example isolates the label construction and next-token shift.
import torch
from torch import nn
from torch.nn import functional as F
torch.manual_seed(42)
IGNORE = -100
rows = [
{
"context": ["system", "inspect", "the", "schema"],
"target": ["inspect_schema"],
},
{
"context": ["system", "run", "a", "safe", "probe"],
"target": ["run_sql_query", "SELECT", "COUNT"],
},
]
vocabulary = {"<pad>": 0}
for row in rows:
for token in row["context"] + row["target"]:
vocabulary.setdefault(token, len(vocabulary))
def encode(row):
context_ids = [vocabulary[token] for token in row["context"]]
target_ids = [vocabulary[token] for token in row["target"]]
input_ids = torch.tensor(context_ids + target_ids, dtype=torch.long)
labels = input_ids.clone()
labels[: len(context_ids)] = IGNORE
return input_ids, labels
examples = [encode(row) for row in rows]
assert all(
labels[: len(row["context"])].eq(IGNORE).all()
for row, (_, labels) in zip(rows, examples)
)
assert all(
labels[len(row["context"]) :].ne(IGNORE).all()
for row, (_, labels) in zip(rows, examples)
)
class TinyCausalLM(nn.Module):
def __init__(self, vocab_size, hidden_size=32):
super().__init__()
self.embedding = nn.Embedding(vocab_size, hidden_size)
self.rnn = nn.GRU(hidden_size, hidden_size, batch_first=True)
self.lm_head = nn.Linear(hidden_size, vocab_size)
def forward(self, input_ids):
hidden, _ = self.rnn(self.embedding(input_ids))
return self.lm_head(hidden)
model = TinyCausalLM(len(vocabulary))
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
for _ in range(100):
for input_ids, labels in examples:
optimizer.zero_grad()
logits = model(input_ids.unsqueeze(0))
shifted_logits = logits[:, :-1, :]
shifted_labels = labels.unsqueeze(0)[:, 1:]
loss = F.cross_entropy(
shifted_logits.reshape(-1, len(vocabulary)),
shifted_labels.reshape(-1),
ignore_index=IGNORE,
)
loss.backward()
optimizer.step()
print("masked SFT loss:", float(loss))
The context tokens are present in input_ids so the model can condition on them, but their entries in labels are -100 and do not contribute to cross-entropy. Shifting both tensors means the logit at position t predicts the label at position t + 1; the first target token is therefore predicted from the final context token. The assertions make the intended assistant span explicit.
In the real pipeline, the teacher first ran the fixed SQL harness and only successful trajectories entered the dataset. Each assistant decision became a row with the preceding conversation and one canonical teacher action as the target. The student then trained with ordinary hard-token SFT on those rows. Evaluation returned the adapter to the same harness, where it had to generate its own actions and handle its own observations.
Dataset And Filtering
The teacher runs inside the same harness the student will be evaluated in: the actual SQL environment, not a detached synthetic dataset.

The data-generation flow was:
- Split train and eval tasks.
- Run the teacher on train tasks.
- Keep only successful trajectories.
- Turn each successful teacher decision into an SFT row.
- Canonicalize the assistant decision target.
- Render the full training sequence.
- Filter rows that do not fit the training context budget.
- Split kept rows into train and validation.

I kept only successful trajectories. That is conservative, but it keeps the dataset’s meaning clean. A failed trajectory can contain locally reasonable actions followed by a wrong final submit. If I kept those earlier actions, I would be claiming I know which pieces of a failed plan deserve credit. For the first experiment, I avoided that ambiguity.
Both teacher-data rounds used the same 879 train tasks. In the next table, success, submitted count, parsed actions, and average turns are measured over all 879 attempted train tasks. Source SFT rows count only assistant decisions from successful trajectories, so they are not simply success multiplied by average turns.
| Teacher source | Success | Submitted | Other stops | Parsed actions | Avg turns / attempted task | Source SFT rows |
|---|---|---|---|---|---|---|
| GPT 5.5 medium | 446/879 = 50.7% | 879 | 0 parse / 0 repeat / 0 max-runtime | 2262 | 2.57 | 1046 |
| Qwen3.5-35B-A3B 8-bit | 394/879 = 44.8% | 816 | 0 parse / 32 repeat / 31 max-runtime | 3064 | 3.49 | 1232 |
The SFT row count is larger than the number of successful tasks because the unit of supervision is one decision point inside a successful trajectory.

The length filter was based on the full rendered training sequence, not only the output length. A row includes the system message, user task, buggy SQL, prior assistant decisions, tool observations, and the target assistant decision. The target decision is usually short. The context can be long because schema text and query observations accumulate across turns.
| Dataset / training path | Source rows | Kept at 4096 | Dropped | Final train / validation rows | In-loop validation | Token length min / P50 / P90 / P95 percentiles |
|---|---|---|---|---|---|---|
| GPT 5.5 rows, Qwen SFT | 1046 | 1042 | 4 | 990 / 52 | enabled | 605 / 1786 / 2948 / 3208 |
| Qwen3.5-35B-A3B 8-bit rows, Qwen SFT | 1232 | 1211 | 21 | 1151 / 60 | enabled | 604 / 2014 / 3180 / 3531 |
| Qwen3.5-35B-A3B 8-bit rows, LFM SFT | 1232 | 1226 | 6 | 1165 / 61 | disabled for LFM | 569 / 1907 / 2984 / 3335 |
The longest source rows before filtering were 15836 tokens for the GPT-row Qwen path, 5908 for the Qwen-row Qwen path, and 5297 for the Qwen-row LFM path. The longer rows were not “bad examples.” They just did not fit the training infrastructure.
The LFM train/validation split was still written, but trainer validation was disabled because validation compilation was unreliable. The 220-task harness eval remained the comparison point.
I also estimated how much of each SFT row was prompt/history versus assistant target.
| SFT row token estimate | Teacher rows | Mean | P50 | P90 | P95 |
|---|---|---|---|---|---|
| Prompt/history before target | GPT 5.5 | 1339 | 1345 | 2367 | 2552 |
| Target action JSON | GPT 5.5 | 90 | 69 | 192 | 259 |
| Prompt plus target | GPT 5.5 | 1431 | 1447 | 2516 | 2745 |
| Prompt/history before target | Qwen3.5-35B-A3B 8-bit | 1558 | 1556 | 2578 | 2862 |
| Target action JSON | Qwen3.5-35B-A3B 8-bit | 83 | 71 | 150 | 191 |
| Prompt plus target | Qwen3.5-35B-A3B 8-bit | 1643 | 1663 | 2693 | 3013 |
This is the shape I expected for agents. The model reads a relatively large state and emits a compact action.
Training Setup
All benchmark results used the same 220-task held-out eval split and the same SQL-agent harness. Teacher traces were generated only on train tasks. I did not train on hidden tests, reference SQL, failed teacher trajectories, or teacher free-form prose.
| Piece | Setup |
|---|---|
| Local workbench | Mac with 128 GB unified memory for dataset prep, MLX serving/eval, GPT evals, charts, and blog artifacts |
| Training machine | Rented NVIDIA RTX 3090 CUDA server for LoRA SFT and direct student evals |
| First teacher | GPT 5.5 medium |
| Same-family teacher | mlx-community/Qwen3.5-35B-A3B-8bit |
| Hosted smaller baseline | GPT 5.4 mini |
| Students | unsloth/Qwen3.5-0.8B, unsloth/Qwen3.5-2B, LiquidAI/LFM2.5-8B-A1B |
The adapters made it practical to run several student experiments quickly on the GPU machine.
| Setting | Value |
|---|---|
| Backend | CUDA / Unsloth-style SFT |
| Epochs | 3 |
| Optimizer updates | 372 for GPT rows; 432 for Qwen-row Qwen students; 438 for Qwen-row LFM |
| Batch / grad accumulation / effective batch | 1 / 8 / 8 |
| Learning rate / seed | 5e-5 / 42 |
| LoRA rank / alpha | 32 / 32 |
| Precision | bf16, 16-bit LoRA, not 4-bit |
| Qwen target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| LFM target modules | in_proj, out_proj, q_proj, k_proj, v_proj, w1, w2, w3 |
| Student | Trainable LoRA parameters |
|---|---|
| Qwen3.5-0.8B | 12.78M of 865.77M, about 1.48% |
| Qwen3.5-2B | 21.82M of 2.24B, about 0.98% |
| LFM2.5-8B-A1B | 11.40M of 8.48B, about 0.13% |
I used this as a boring first recipe, not a hyperparameter search; the fixed settings kept the runs comparable. I avoided 4-bit training because the models fit with 16-bit LoRA on the GPU path, and I wanted the first result to be less entangled with quantization.
The local student eval cap includes both the input/history tokens and the generated action tokens. Max new tokens is still shown separately because it controls the maximum size of one assistant action.
| Run family | Context / sequence cap | Output budget per turn | Max turns | Timeout | Temperature |
|---|---|---|---|---|---|
| CUDA local Hugging Face PEFT students and bases | 8192-token total sequence cap | 512 new tokens | 8 | 180s/task | 0.0 |
| GPT 5.5 medium teacher eval | model/provider context | 1024 new tokens | 8 | 180s/task | 0.0 |
| GPT 5.4 mini hosted eval | model/provider context | 2048 new tokens | 8 | 180s/task | 0.0 |
| Qwen3.5-35B-A3B 8-bit MLX eval | MLX model context | 2048 new tokens | 8 | 180s/task | 0.0 |
| LiquidAI/LFM2.5-8B-A1B-MLX-8bit base eval | MLX model context | 2048 new tokens | 8 | 180s/task | 0.0 |
Results As Research Questions
I find the results easier to read as questions. The main question is about the 0.8B student, but I also include Qwen3.5-2B and LiquidAI/LFM2.5-8B-A1B as comparison students. They help answer whether the behavior is a tiny-model accident or a broader effect of the training setup.
A base row means the model ran in the harness without a LoRA adapter. An SFT row means the base model plus the trained adapter ran in the same harness. Success means the final submitted SQL passed the hidden tests.
First: how good are the strong baselines on the same harness?

That gives the scale. Qwen3.5-35B-A3B 8-bit solved 96/220, GPT 5.4 mini solved 105/220, and GPT 5.5 medium solved 115/220. The GPT baselines submitted on every task and had no parse, repeat, max-turn, or runtime stops. Their failures were almost all wrong SQL submissions, which is a healthier failure shape than failing to operate the loop at all.
Second: what does hard-token trajectory SFT do to small students?

It teaches them to act. Qwen3.5-0.8B moved from 1/220 to 44/220 on GPT 5.5 medium rows, and to 46/220 on Qwen3.5-35B-A3B 8-bit rows. Qwen3.5-2B moved from 0/220 to 57/220 and 58/220. LiquidAI/LFM2.5-8B-A1B-MLX-8bit started higher at 9/220, then reached 47/220 with GPT rows and 50/220 with Qwen rows.
The important behavior change is submission. The base Qwen models submitted only once. After SFT, they submitted on most tasks. That means the training transferred the protocol: inspect, query, read, submit. But most remaining failures were wrong submissions, so SQL judgment did not transfer at the same level.
Third: does a same-family Qwen teacher help?

The same-family teacher helped slightly on success: +2, +1, and +3 tasks for the three students. The surprise is the cost. Repeat stops rose from 10 to 54 for Qwen3.5-0.8B, from 9 to 65 for Qwen3.5-2B, and from 27 to 38 for LiquidAI/LFM2.5-8B-A1B.
In that chart, the left side is solved eval tasks, where higher is better. The right side is repeated-action stops, where lower is better. I think of that right side as the loop-control cost. Qwen3.5-35B-A3B 8-bit produced more SFT rows because its successful trajectories were longer. Those rows gave small success gains, but they also transferred a loopier policy.
Fourth: what changed in the harness behavior?

This is the chart that explains the experiment. The base Qwen students are mostly repeat-stop bars; they are not consistently reaching the final answer step. GPT-row SFT turns those bars into mostly submissions.
The cleanest summary is:
SFT changed the students from “does not really operate the harness” into “operates the harness, but often submits wrong SQL.”
Exact Numbers Behind The Charts
The table below gives the values behind the charts. Success, wrong submits, parse stops, repeat stops, and max/runtime stops are mutually exclusive final task outcomes, so they sum to 220. SQL execution errors are diagnostic events inside a run and can overlap with any of those outcomes.
| Strong baseline | Success | Wrong submits | Parse stops | Repeat stops | Max/runtime stops | SQL-error tasks (non-exclusive) |
|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3B 8-bit local MLX eval | 96/220 = 43.6% | 106 | 0 | 9 | 9 | 4 |
| GPT 5.4 mini hosted baseline | 105/220 = 47.7% | 115 | 0 | 0 | 0 | 4 |
| GPT 5.5 medium teacher/baseline | 115/220 = 52.3% | 105 | 0 | 0 | 0 | 1 |
| Student run | Success | Wrong submits | Parse stops | Repeat stops | Max/runtime stops | SQL-error tasks (non-exclusive) | Avg turns |
|---|---|---|---|---|---|---|---|
| Qwen3.5-0.8B base | 1/220 = 0.5% | 0 | 11 | 208 | 0 | 0 | 2.39 |
| Qwen3.5-0.8B SFT on GPT 5.5 rows | 44/220 = 20.0% | 160 | 6 | 10 | 0 | 45 | 2.50 |
| Qwen3.5-0.8B SFT on Qwen3.5-35B-A3B 8-bit rows | 46/220 = 20.9% | 109 | 7 | 54 | 4 | 5 | 3.48 |
| Qwen3.5-2B base | 0/220 = 0.0% | 1 | 0 | 219 | 0 | 0 | 2.00 |
| Qwen3.5-2B SFT on GPT 5.5 rows | 57/220 = 25.9% | 149 | 5 | 9 | 0 | 49 | 2.59 |
| Qwen3.5-2B SFT on Qwen3.5-35B-A3B 8-bit rows | 58/220 = 26.4% | 91 | 3 | 65 | 3 | 2 | 3.60 |
| LiquidAI/LFM2.5-8B-A1B-MLX-8bit base | 9/220 = 4.1% | 22 | 0 | 52 | 137 | 8 | 1.64 |
| LiquidAI/LFM2.5-8B-A1B SFT on GPT 5.5 rows | 47/220 = 21.4% | 139 | 7 | 27 | 0 | 34 | 2.59 |
| LiquidAI/LFM2.5-8B-A1B SFT on Qwen3.5-35B-A3B 8-bit rows | 50/220 = 22.7% | 120 | 6 | 38 | 6 | 12 | 3.52 |
Failure Analysis
Success rate alone hides too much for agents. I also looked at turns, generated tokens, prompt estimates, validation loss, and stop reasons.

The student generated-token chart makes the same-family teacher effect visible. Students trained on Qwen3.5-35B-A3B 8-bit rows did not mainly become more verbose per action; their actions stayed short structured JSON, but they took more turns. More turns means more generated output per task and more chances to drift into a state the teacher never supervised.
The LiquidAI/LFM2.5-8B-A1B-MLX-8bit base row is approximate because many invalid BAML generations were not saved as clean assistant messages. I used saved output-token counters for invalid generations and saved parsed outputs where available. The pattern is still useful: the base LFM output was raw and verbose, while fine-tuned LFM output was closer to short executable actions.

The teacher and strong-baseline token chart uses its own scale. Qwen3.5-35B-A3B 8-bit often produced shorter action text per call than GPT 5.5 medium, but it took more turns, so its generated output per task ended up higher. GPT 5.4 mini was the best efficiency point among the strong baselines I measured: close to GPT 5.5 medium in task success, with fewer turns and fewer generated tokens per task.
The prompt side is usually larger than the generated side. Each extra turn adds one assistant action, but it also forces the next prompt to resend the system instructions, task, prior actions, schema observations, query results, and structured-output instructions. That is why turns matter twice.
For the initial GPT-row and baseline runs where prompt reconstruction was available, the prompt and total estimates looked like this:
| Run | Prompt/call mean | Prompt/call P90 | Total/task mean |
|---|---|---|---|
| Qwen3.5-0.8B base | 1118 | 1993 | 2881 |
| Qwen3.5-0.8B SFT, GPT 5.5 medium rows | 1568 | 2545 | 4130 |
| Qwen3.5-2B base | 1399 | 2420 | 2878 |
| Qwen3.5-2B SFT, GPT 5.5 medium rows | 1611 | 2586 | 4400 |
| LiquidAI/LFM2.5-8B-A1B SFT, GPT 5.5 medium rows | 1620 | 2600 | 4412 |
| Qwen3.5-35B-A3B 8-bit local eval | 1872 | 2939 | 7000 |
| GPT 5.4 mini hosted baseline | 1471 | 2522 | 3354 |
| GPT 5.5 medium teacher/baseline | 1590 | 2601 | 4287 |
Validation loss shows the teacher-forcing gap. For Qwen3.5-0.8B, final validation loss improved from 0.3494 on GPT 5.5 rows to 0.2553 on Qwen3.5-35B-A3B 8-bit rows. For Qwen3.5-2B, it improved from 0.2966 to 0.2051. Under teacher-forced SFT, the same-family rows were clearly easier for the Qwen students to predict.
The rollout improvement was much smaller. Validation loss asks, “Can the model predict the teacher’s next action given a teacher-style context?” The harness asks, “Can the model recover from its own previous actions, choose useful SQL probes, and submit SQL that passes hidden tests?” Those are related questions, but they are not the same measurement.
Hardware And Infrastructure Lessons
The Qwen3.5-35B-A3B 8-bit teacher-row runs took longer because they kept more rows, ran more optimizer updates, and had longer median context.

Wall time moved from 15.9m to 22.4m for Qwen3.5-0.8B, from 22.1m to 29.4m for Qwen3.5-2B, and from 57.4m to 71.8m for LiquidAI/LFM2.5-8B-A1B.
The Mac with 128 GB unified memory was useful for capacity-heavy local work: preparing data, generating charts, and serving quantized models. But training is repeated forward, backward, and optimizer math, where CUDA kernels, tensor cores, FlashAttention, and GPU-local VRAM matter more than total host memory. That is the RAM versus VRAM lesson I learned the slow way: a 24 GB VRAM GPU can be much faster for adapter training than a 128 GB unified-memory Mac, even though the Mac has more total memory.
LiquidAI/LFM2.5-8B-A1B took the longest because it is a much larger model and used an eager expert-execution path in this training setup.
What I Learned
The first lesson is that trajectory SFT can transfer the loop: after SFT, the students inspected, queried, and submitted far more often.
The second lesson is that the failure simply moved downstream. A wrong submit is better than no submit, but it is still a failed task. The students learned the protocol faster than they learned SQL judgment under their own rollouts.
The third lesson is that successful-only filtering made the first dataset easier to reason about. I probably threw away some locally good actions from failed trajectories, but I avoided teaching from traces whose credit assignment I could not defend.
The fourth lesson is that a same-family teacher can be easier to imitate and still transfer bad habits: slightly better success and lower validation loss, but more repeated-action behavior.
The fifth lesson is that teacher-forced loss and rollout success measure different things: validation loss dropped sharply on same-family rows while the eval gain was tiny and loop control got worse. I learned not to trust offline loss alone as a proxy for agent quality.
Conclusion
Can a 0.8B model learn to act like a stronger model inside a real tool-use loop?
Yes, but only in the specific sense this experiment measured. For the 0.8B student itself, offline teacher-trace hard-token SFT taught the agent protocol: unsloth/Qwen3.5-0.8B went from 1/220 to 44/220 on GPT 5.5 medium teacher rows and 46/220 on Qwen3.5-35B-A3B 8-bit rows.
The larger student runs are comparison points, not the answer to the 0.8B question. They show the same broad pattern: SFT teaches the loop, but it does not close the gap to the stronger teachers.
But the students did not inherit teacher-level SQL judgment; the stronger baselines remained far ahead. That is the honest result of this first post: protocol transfer, yes. Teacher-level decision making, no.
The natural next step is to keep the traces and the harness fixed but give the student a richer signal than one hard token per position. That is what the next post does: off-policy top-k soft-label distillation over these same teacher traces.