← ResearchSmall-Model Distillation

Small-Model Distillation, Part 1: Offline Teacher-Trace SFT for a 0.8B SQL Agent

Isaac Kargar25 min read

TL;DR

I wanted to see whether a very small language model could learn to act inside a real SQL tool-use loop. The model was not asked to only write a final answer. It had to inspect a SQLite database, run SQL probes, read tool observations, and submit corrected SQL that passed hidden deterministic tests.

The first method in this series is the simplest useful one: offline teacher-trace hard-token SFT. A stronger teacher runs the fixed harness first. I keep only successful train trajectories. Then each teacher decision becomes a supervised fine-tuning (SFT) row of the form conversation so far -> next structured decision.

RoleResult
Main 0.8B studentunsloth/Qwen3.5-0.8B moved from 1/220 to 44/220 with GPT 5.5 medium rows, and to 46/220 with Qwen3.5-35B-A3B 8-bit rows.
Comparison studentsunsloth/Qwen3.5-2B reached 58/220 at best; LiquidAI/LFM2.5-8B-A1B reached 50/220 at best.
Strong baselinesQwen3.5-35B-A3B 8-bit solved 96/220, GPT 5.4 mini solved 105/220, and GPT 5.5 medium solved 115/220.

The answer is not “the small model became as good as the teacher.” It did not. The better answer is: hard-token trajectory SFT transferred the agent protocol, but not teacher-level SQL judgment. The students learned to operate the harness and submit much more often, but most of the remaining failures moved downstream into wrong SQL submissions and loop-control mistakes.

The Method

The reusable idea in this post is a hard-token SFT contract. The SQL-agent code runs the teacher in the harness, keeps successful trajectories, and turns each assistant decision into a messages -> assistant target row. Training masks the context and learns only the teacher-written hard target with LoRA adapters. The complete teaching example below shows the label construction and next-token shift.

What I Wanted To Test

I am interested in small language models that become good at a narrow, useful task. Not a general chat model. Not a leaderboard toy. A small model that can sit inside a real workflow and do one thing well enough to matter.

For this first post, I started with a SQL repair agent. Each task gives the model a user issue, a buggy SQL query, and a SQLite database. The model can inspect the schema, run SQL queries, read observations, and eventually submit corrected SQL. The final score is deterministic: the submitted SQL either passes hidden tests or it does not.

The question was:

Can a tiny model learn the behavior of a stronger SQL tool-use agent from saved teacher trajectories?

That wording matters. I was not only asking whether the model can write SQL-looking text. I wanted the model to learn the loop: choose an action, respect the schema, read the observation, decide what to do next, and stop when it has enough evidence.

I kept the benchmark, action schema, parser, stop rules, and eval split fixed. Small-agent results are easy to distort by changing the harness, so I did not tune the environment per model.

The SQL-Agent Problem

Here is the kind of task the model sees:

Database: chinook

User issue: I want to find the latest track_id and use that id to filter records in the track table.

Buggy SQL:

WITH vars AS (SELECT COUNT(*) AS vars_id FROM track)
SELECT * FROM track WHERE track_id = vars_id

The bug is small but meaningful. COUNT(*) is not the latest id. The intended repair is closer to MAX(track_id). The buggy SQL gives useful clues about tables and columns, but it can also anchor the model to the wrong operation.

That gives the model two jobs:

  • Operate the agent protocol correctly.
  • Choose the right SQL repair.

A weak model can fail before doing SQL reasoning at all. It can inspect schema, inspect schema again, emit invalid JSON, or repeat an unproductive action until the harness stops it. A better model can drive the harness but still submit SQL that fails the hidden tests. Those are different failure modes, and I wanted the experiment to separate them.

A SQL-agent task flows from a user request and buggy query through schema inspection and SQL probes to a submitted query checked against hidden tests
A task supplies the user issue, buggy SQL, database, and hidden tests; the model sees the issue, query, and tool observations.

The benchmark source was birdsql/six-gym-sqlite.

Concretely, the benchmark filter was category == "Query" and db_id in {"netflix", "movie_3", "books", "chinook"}. I chose this fixed slice to keep the first experiment narrow while preserving several domains, then split inside each database so train and eval kept the same domain mix.

SettingValue
Task categoryQuery
Databasesnetflix, movie_3, books, chinook
Source rows scanned5000
Candidate rows after filtering1099
Train split879 tasks
Eval split220 tasks
Split seed42

The split by database was:

DatabaseCandidate tasksTrain tasksEval tasks
books28222656
chinook25120150
movie_327321855
netflix29323459
Total1099879220

The model can choose only three structured actions:

{"action": "inspect_schema"}
{"action": "run_sql_query", "sql": "SELECT ..."}
{"action": "submit_sql", "sql": ["SQL statement 1", "SQL statement 2"]}

BAML, or Boundary Markup Language, is the structured-output layer around the prompt contract. It defines the action schema and renders the model request. Hosted teacher and baseline calls use the BAML structured parsing path; local HF/PEFT student eval decodes model text and the harness validates it with parse_decision. In both paths, the accepted target is canonical decision JSON, not free-form teacher prose.

The loop is simple:

  1. Build messages from the task.
  2. Ask the model for one structured action.
  3. Parse and validate the action.
  4. Execute the action in SQLite when needed.
  5. Append the observation.
  6. Repeat until a stop condition below.

The stop categories are part of the result, not just logging noise.

Stop reasonWhat it means
submittedThe model submitted final SQL; the SQL either passed or failed hidden tests
parse_failureThe model did not produce a valid structured action
repeated_actionThe model repeated an action that the harness considered unproductive
max_turnsThe model kept acting but never reached a valid final submission
runtime_errorThe model call or harness call failed during the task

For this task, formatting is not cosmetic. A human can understand “I should inspect the schema first,” but the harness needs {"action":"inspect_schema"}. I did not use keyword matching or an LLM judge to rescue malformed actions.

The Training Idea

This post is about post-training: shaping a base model’s behavior for a use case, not pretraining.

Teacher trajectories are collected in the fixed harness, converted into assistant-decision rows, and used for student SFT
This post focuses on post-training task distillation from verified teacher trajectories.

Knowledge distillation means using a stronger teacher model to transfer useful behavior into a smaller student model. For agents, the important question is not only “which teacher?” It is also:

Which part of the teacher behavior becomes supervision?

Distillation signalWhat the student learns fromWhy it matters
Hard labelsThe teacher’s chosen output tokensSimple SFT: given this input, produce this exact target.
Soft labels / logitsThe teacher’s probability distribution over next tokensPreserves uncertainty over alternatives instead of only the winning token.
Feature distillationInternal teacher activations or hidden statesTries to align representations, but needs direct teacher internals.
Final-answer distillationOnly the completed teacher answerUseful for answer-only tasks, weak for agents because it discards the process.
Trajectory distillationIntermediate actions, observations, and final answerTeaches the policy: inspect, query, react, and submit.
Reward or RL-style methodsA scalar outcome after rolloutOptimizes behavior through environment feedback instead of direct imitation.
Hard targets, soft token probabilities, intermediate features, trajectories, and rewards provide different distillation signals
Hard tokens, soft probabilities, features, trajectories, and rewards provide different kinds of distillation supervision.

This first post uses offline hard-token trajectory distillation. Offline means the teacher runs first and training happens later from saved traces. Hard-token means the student trains on the exact tokens the teacher chose, not the teacher’s probability distribution. Trajectory means the training rows come from intermediate tool-use decisions, not only final SQL.

So the unit of learning is:

Full conversation state before the teacher decision -> canonical next executable decision.

This is still a local imitation objective. It does not directly optimize task success under the student’s own rollouts. That limitation becomes important in the results.

A Small Educational Code Walkthrough

The runnable example below is a small PyTorch version of the training contract. It uses whitespace tokens and a tiny causal model so it can run without downloading a language model. The actual experiment used Qwen’s chat template, canonical BAML decisions, and LoRA adapters; this example isolates the label construction and next-token shift.

import torch
from torch import nn
from torch.nn import functional as F

torch.manual_seed(42)
IGNORE = -100

rows = [
    {
        "context": ["system", "inspect", "the", "schema"],
        "target": ["inspect_schema"],
    },
    {
        "context": ["system", "run", "a", "safe", "probe"],
        "target": ["run_sql_query", "SELECT", "COUNT"],
    },
]

vocabulary = {"<pad>": 0}
for row in rows:
    for token in row["context"] + row["target"]:
        vocabulary.setdefault(token, len(vocabulary))


def encode(row):
    context_ids = [vocabulary[token] for token in row["context"]]
    target_ids = [vocabulary[token] for token in row["target"]]
    input_ids = torch.tensor(context_ids + target_ids, dtype=torch.long)
    labels = input_ids.clone()
    labels[: len(context_ids)] = IGNORE
    return input_ids, labels


examples = [encode(row) for row in rows]
assert all(
    labels[: len(row["context"])].eq(IGNORE).all()
    for row, (_, labels) in zip(rows, examples)
)
assert all(
    labels[len(row["context"]) :].ne(IGNORE).all()
    for row, (_, labels) in zip(rows, examples)
)


class TinyCausalLM(nn.Module):
    def __init__(self, vocab_size, hidden_size=32):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, hidden_size)
        self.rnn = nn.GRU(hidden_size, hidden_size, batch_first=True)
        self.lm_head = nn.Linear(hidden_size, vocab_size)

    def forward(self, input_ids):
        hidden, _ = self.rnn(self.embedding(input_ids))
        return self.lm_head(hidden)


model = TinyCausalLM(len(vocabulary))
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)

for _ in range(100):
    for input_ids, labels in examples:
        optimizer.zero_grad()
        logits = model(input_ids.unsqueeze(0))
        shifted_logits = logits[:, :-1, :]
        shifted_labels = labels.unsqueeze(0)[:, 1:]
        loss = F.cross_entropy(
            shifted_logits.reshape(-1, len(vocabulary)),
            shifted_labels.reshape(-1),
            ignore_index=IGNORE,
        )
        loss.backward()
        optimizer.step()

print("masked SFT loss:", float(loss))

The context tokens are present in input_ids so the model can condition on them, but their entries in labels are -100 and do not contribute to cross-entropy. Shifting both tensors means the logit at position t predicts the label at position t + 1; the first target token is therefore predicted from the final context token. The assertions make the intended assistant span explicit.

In the real pipeline, the teacher first ran the fixed SQL harness and only successful trajectories entered the dataset. Each assistant decision became a row with the preceding conversation and one canonical teacher action as the target. The student then trained with ordinary hard-token SFT on those rows. Evaluation returned the adapter to the same harness, where it had to generate its own actions and handle its own observations.

Dataset And Filtering

The teacher runs inside the same harness the student will be evaluated in: the actual SQL environment, not a detached synthetic dataset.

Successful teacher episodes in the SQL harness are filtered and split into one supervised row for each assistant decision
Successful teacher episodes are replayed in the SQL harness, and each verified assistant action becomes an SFT row.

The data-generation flow was:

  1. Split train and eval tasks.
  2. Run the teacher on train tasks.
  3. Keep only successful trajectories.
  4. Turn each successful teacher decision into an SFT row.
  5. Canonicalize the assistant decision target.
  6. Render the full training sequence.
  7. Filter rows that do not fit the training context budget.
  8. Split kept rows into train and validation.
One verified teacher trajectory becomes sequential SFT rows, each pairing the preceding state with the next assistant action
One successful trajectory yields sequential rows for the next action after each tool observation.

I kept only successful trajectories. That is conservative, but it keeps the dataset’s meaning clean. A failed trajectory can contain locally reasonable actions followed by a wrong final submit. If I kept those earlier actions, I would be claiming I know which pieces of a failed plan deserve credit. For the first experiment, I avoided that ambiguity.

Both teacher-data rounds used the same 879 train tasks. In the next table, success, submitted count, parsed actions, and average turns are measured over all 879 attempted train tasks. Source SFT rows count only assistant decisions from successful trajectories, so they are not simply success multiplied by average turns.

Teacher sourceSuccessSubmittedOther stopsParsed actionsAvg turns / attempted taskSource SFT rows
GPT 5.5 medium446/879 = 50.7%8790 parse / 0 repeat / 0 max-runtime22622.571046
Qwen3.5-35B-A3B 8-bit394/879 = 44.8%8160 parse / 32 repeat / 31 max-runtime30643.491232

The SFT row count is larger than the number of successful tasks because the unit of supervision is one decision point inside a successful trajectory.

The data pipeline splits tasks, runs the teacher, keeps successful trajectories, renders rows, filters long sequences, and makes train and validation sets
The data pipeline splits tasks, filters successful teacher traces by context length, and creates the SFT dataset.

The length filter was based on the full rendered training sequence, not only the output length. A row includes the system message, user task, buggy SQL, prior assistant decisions, tool observations, and the target assistant decision. The target decision is usually short. The context can be long because schema text and query observations accumulate across turns.

Dataset / training pathSource rowsKept at 4096DroppedFinal train / validation rowsIn-loop validationToken length min / P50 / P90 / P95 percentiles
GPT 5.5 rows, Qwen SFT104610424990 / 52enabled605 / 1786 / 2948 / 3208
Qwen3.5-35B-A3B 8-bit rows, Qwen SFT12321211211151 / 60enabled604 / 2014 / 3180 / 3531
Qwen3.5-35B-A3B 8-bit rows, LFM SFT1232122661165 / 61disabled for LFM569 / 1907 / 2984 / 3335

The longest source rows before filtering were 15836 tokens for the GPT-row Qwen path, 5908 for the Qwen-row Qwen path, and 5297 for the Qwen-row LFM path. The longer rows were not “bad examples.” They just did not fit the training infrastructure.

The LFM train/validation split was still written, but trainer validation was disabled because validation compilation was unreliable. The 220-task harness eval remained the comparison point.

I also estimated how much of each SFT row was prompt/history versus assistant target.

SFT row token estimateTeacher rowsMeanP50P90P95
Prompt/history before targetGPT 5.51339134523672552
Target action JSONGPT 5.59069192259
Prompt plus targetGPT 5.51431144725162745
Prompt/history before targetQwen3.5-35B-A3B 8-bit1558155625782862
Target action JSONQwen3.5-35B-A3B 8-bit8371150191
Prompt plus targetQwen3.5-35B-A3B 8-bit1643166326933013

This is the shape I expected for agents. The model reads a relatively large state and emits a compact action.

Training Setup

All benchmark results used the same 220-task held-out eval split and the same SQL-agent harness. Teacher traces were generated only on train tasks. I did not train on hidden tests, reference SQL, failed teacher trajectories, or teacher free-form prose.

PieceSetup
Local workbenchMac with 128 GB unified memory for dataset prep, MLX serving/eval, GPT evals, charts, and blog artifacts
Training machineRented NVIDIA RTX 3090 CUDA server for LoRA SFT and direct student evals
First teacherGPT 5.5 medium
Same-family teachermlx-community/Qwen3.5-35B-A3B-8bit
Hosted smaller baselineGPT 5.4 mini
Studentsunsloth/Qwen3.5-0.8B, unsloth/Qwen3.5-2B, LiquidAI/LFM2.5-8B-A1B

The adapters made it practical to run several student experiments quickly on the GPU machine.

SettingValue
BackendCUDA / Unsloth-style SFT
Epochs3
Optimizer updates372 for GPT rows; 432 for Qwen-row Qwen students; 438 for Qwen-row LFM
Batch / grad accumulation / effective batch1 / 8 / 8
Learning rate / seed5e-5 / 42
LoRA rank / alpha32 / 32
Precisionbf16, 16-bit LoRA, not 4-bit
Qwen target modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
LFM target modulesin_proj, out_proj, q_proj, k_proj, v_proj, w1, w2, w3
StudentTrainable LoRA parameters
Qwen3.5-0.8B12.78M of 865.77M, about 1.48%
Qwen3.5-2B21.82M of 2.24B, about 0.98%
LFM2.5-8B-A1B11.40M of 8.48B, about 0.13%

I used this as a boring first recipe, not a hyperparameter search; the fixed settings kept the runs comparable. I avoided 4-bit training because the models fit with 16-bit LoRA on the GPU path, and I wanted the first result to be less entangled with quantization.

The local student eval cap includes both the input/history tokens and the generated action tokens. Max new tokens is still shown separately because it controls the maximum size of one assistant action.

Run familyContext / sequence capOutput budget per turnMax turnsTimeoutTemperature
CUDA local Hugging Face PEFT students and bases8192-token total sequence cap512 new tokens8180s/task0.0
GPT 5.5 medium teacher evalmodel/provider context1024 new tokens8180s/task0.0
GPT 5.4 mini hosted evalmodel/provider context2048 new tokens8180s/task0.0
Qwen3.5-35B-A3B 8-bit MLX evalMLX model context2048 new tokens8180s/task0.0
LiquidAI/LFM2.5-8B-A1B-MLX-8bit base evalMLX model context2048 new tokens8180s/task0.0

Results As Research Questions

I find the results easier to read as questions. The main question is about the 0.8B student, but I also include Qwen3.5-2B and LiquidAI/LFM2.5-8B-A1B as comparison students. They help answer whether the behavior is a tiny-model accident or a broader effect of the training setup.

A base row means the model ran in the harness without a LoRA adapter. An SFT row means the base model plus the trained adapter ran in the same harness. Success means the final submitted SQL passed the hidden tests.

First: how good are the strong baselines on the same harness?

Solved-task counts for the Qwen and GPT teacher baselines on the 220-task evaluation
The chart compares solved-task counts for the Qwen and GPT baselines on the fixed 220-task evaluation.

That gives the scale. Qwen3.5-35B-A3B 8-bit solved 96/220, GPT 5.4 mini solved 105/220, and GPT 5.5 medium solved 115/220. The GPT baselines submitted on every task and had no parse, repeat, max-turn, or runtime stops. Their failures were almost all wrong SQL submissions, which is a healthier failure shape than failing to operate the loop at all.

Second: what does hard-token trajectory SFT do to small students?

Solved-task counts for the 0.8B, 2B, and 8B students before and after trajectory SFT
Trajectory SFT raises solved-task counts for the 0.8B, 2B, and 8B students on the same evaluation.

It teaches them to act. Qwen3.5-0.8B moved from 1/220 to 44/220 on GPT 5.5 medium rows, and to 46/220 on Qwen3.5-35B-A3B 8-bit rows. Qwen3.5-2B moved from 0/220 to 57/220 and 58/220. LiquidAI/LFM2.5-8B-A1B-MLX-8bit started higher at 9/220, then reached 47/220 with GPT rows and 50/220 with Qwen rows.

The important behavior change is submission. The base Qwen models submitted only once. After SFT, they submitted on most tasks. That means the training transferred the protocol: inspect, query, read, submit. But most remaining failures were wrong submissions, so SQL judgment did not transfer at the same level.

Third: does a same-family Qwen teacher help?

Success and repeated-action counts after training on GPT versus same-family Qwen traces
The two panels compare success and repeated-action stops after training on GPT rows or same-family Qwen rows.

The same-family teacher helped slightly on success: +2, +1, and +3 tasks for the three students. The surprise is the cost. Repeat stops rose from 10 to 54 for Qwen3.5-0.8B, from 9 to 65 for Qwen3.5-2B, and from 27 to 38 for LiquidAI/LFM2.5-8B-A1B.

In that chart, the left side is solved eval tasks, where higher is better. The right side is repeated-action stops, where lower is better. I think of that right side as the loop-control cost. Qwen3.5-35B-A3B 8-bit produced more SFT rows because its successful trajectories were longer. Those rows gave small success gains, but they also transferred a loopier policy.

Fourth: what changed in the harness behavior?

SQL-agent stop outcomes before and after hard-token trajectory SFT
The stacked bars show terminal SQL-agent outcomes for base models, SFT students, and strong baselines.

This is the chart that explains the experiment. The base Qwen students are mostly repeat-stop bars; they are not consistently reaching the final answer step. GPT-row SFT turns those bars into mostly submissions.

The cleanest summary is:

SFT changed the students from “does not really operate the harness” into “operates the harness, but often submits wrong SQL.”

Exact Numbers Behind The Charts

The table below gives the values behind the charts. Success, wrong submits, parse stops, repeat stops, and max/runtime stops are mutually exclusive final task outcomes, so they sum to 220. SQL execution errors are diagnostic events inside a run and can overlap with any of those outcomes.

Strong baselineSuccessWrong submitsParse stopsRepeat stopsMax/runtime stopsSQL-error tasks (non-exclusive)
Qwen3.5-35B-A3B 8-bit local MLX eval96/220 = 43.6%1060994
GPT 5.4 mini hosted baseline105/220 = 47.7%1150004
GPT 5.5 medium teacher/baseline115/220 = 52.3%1050001
Student runSuccessWrong submitsParse stopsRepeat stopsMax/runtime stopsSQL-error tasks (non-exclusive)Avg turns
Qwen3.5-0.8B base1/220 = 0.5%011208002.39
Qwen3.5-0.8B SFT on GPT 5.5 rows44/220 = 20.0%1606100452.50
Qwen3.5-0.8B SFT on Qwen3.5-35B-A3B 8-bit rows46/220 = 20.9%109754453.48
Qwen3.5-2B base0/220 = 0.0%10219002.00
Qwen3.5-2B SFT on GPT 5.5 rows57/220 = 25.9%149590492.59
Qwen3.5-2B SFT on Qwen3.5-35B-A3B 8-bit rows58/220 = 26.4%91365323.60
LiquidAI/LFM2.5-8B-A1B-MLX-8bit base9/220 = 4.1%2205213781.64
LiquidAI/LFM2.5-8B-A1B SFT on GPT 5.5 rows47/220 = 21.4%1397270342.59
LiquidAI/LFM2.5-8B-A1B SFT on Qwen3.5-35B-A3B 8-bit rows50/220 = 22.7%1206386123.52

Failure Analysis

Success rate alone hides too much for agents. I also looked at turns, generated tokens, prompt estimates, validation loss, and stop reasons.

Generated-token usage for base and SFT students during SQL-agent evaluation
Estimated generated output tokens and average turns for the student evaluations show the cost of longer agent episodes.

The student generated-token chart makes the same-family teacher effect visible. Students trained on Qwen3.5-35B-A3B 8-bit rows did not mainly become more verbose per action; their actions stayed short structured JSON, but they took more turns. More turns means more generated output per task and more chances to drift into a state the teacher never supervised.

The LiquidAI/LFM2.5-8B-A1B-MLX-8bit base row is approximate because many invalid BAML generations were not saved as clean assistant messages. I used saved output-token counters for invalid generations and saved parsed outputs where available. The pattern is still useful: the base LFM output was raw and verbose, while fine-tuned LFM output was closer to short executable actions.

Generated-token usage for the teacher and strong SQL-agent baselines
Estimated generated output tokens and average turns for the teacher and strong baseline evaluations.

The teacher and strong-baseline token chart uses its own scale. Qwen3.5-35B-A3B 8-bit often produced shorter action text per call than GPT 5.5 medium, but it took more turns, so its generated output per task ended up higher. GPT 5.4 mini was the best efficiency point among the strong baselines I measured: close to GPT 5.5 medium in task success, with fewer turns and fewer generated tokens per task.

The prompt side is usually larger than the generated side. Each extra turn adds one assistant action, but it also forces the next prompt to resend the system instructions, task, prior actions, schema observations, query results, and structured-output instructions. That is why turns matter twice.

For the initial GPT-row and baseline runs where prompt reconstruction was available, the prompt and total estimates looked like this:

RunPrompt/call meanPrompt/call P90Total/task mean
Qwen3.5-0.8B base111819932881
Qwen3.5-0.8B SFT, GPT 5.5 medium rows156825454130
Qwen3.5-2B base139924202878
Qwen3.5-2B SFT, GPT 5.5 medium rows161125864400
LiquidAI/LFM2.5-8B-A1B SFT, GPT 5.5 medium rows162026004412
Qwen3.5-35B-A3B 8-bit local eval187229397000
GPT 5.4 mini hosted baseline147125223354
GPT 5.5 medium teacher/baseline159026014287

Validation loss shows the teacher-forcing gap. For Qwen3.5-0.8B, final validation loss improved from 0.3494 on GPT 5.5 rows to 0.2553 on Qwen3.5-35B-A3B 8-bit rows. For Qwen3.5-2B, it improved from 0.2966 to 0.2051. Under teacher-forced SFT, the same-family rows were clearly easier for the Qwen students to predict.

The rollout improvement was much smaller. Validation loss asks, “Can the model predict the teacher’s next action given a teacher-style context?” The harness asks, “Can the model recover from its own previous actions, choose useful SQL probes, and submit SQL that passes hidden tests?” Those are related questions, but they are not the same measurement.

Hardware And Infrastructure Lessons

The Qwen3.5-35B-A3B 8-bit teacher-row runs took longer because they kept more rows, ran more optimizer updates, and had longer median context.

Training wall time for Qwen and LiquidAI students with each teacher's rows
Measured CUDA LoRA training time in minutes for each student and teacher-row source.

Wall time moved from 15.9m to 22.4m for Qwen3.5-0.8B, from 22.1m to 29.4m for Qwen3.5-2B, and from 57.4m to 71.8m for LiquidAI/LFM2.5-8B-A1B.

The Mac with 128 GB unified memory was useful for capacity-heavy local work: preparing data, generating charts, and serving quantized models. But training is repeated forward, backward, and optimizer math, where CUDA kernels, tensor cores, FlashAttention, and GPU-local VRAM matter more than total host memory. That is the RAM versus VRAM lesson I learned the slow way: a 24 GB VRAM GPU can be much faster for adapter training than a 128 GB unified-memory Mac, even though the Mac has more total memory.

LiquidAI/LFM2.5-8B-A1B took the longest because it is a much larger model and used an eager expert-execution path in this training setup.

What I Learned

The first lesson is that trajectory SFT can transfer the loop: after SFT, the students inspected, queried, and submitted far more often.

The second lesson is that the failure simply moved downstream. A wrong submit is better than no submit, but it is still a failed task. The students learned the protocol faster than they learned SQL judgment under their own rollouts.

The third lesson is that successful-only filtering made the first dataset easier to reason about. I probably threw away some locally good actions from failed trajectories, but I avoided teaching from traces whose credit assignment I could not defend.

The fourth lesson is that a same-family teacher can be easier to imitate and still transfer bad habits: slightly better success and lower validation loss, but more repeated-action behavior.

The fifth lesson is that teacher-forced loss and rollout success measure different things: validation loss dropped sharply on same-family rows while the eval gain was tiny and loop control got worse. I learned not to trust offline loss alone as a proxy for agent quality.

Conclusion

Can a 0.8B model learn to act like a stronger model inside a real tool-use loop?

Yes, but only in the specific sense this experiment measured. For the 0.8B student itself, offline teacher-trace hard-token SFT taught the agent protocol: unsloth/Qwen3.5-0.8B went from 1/220 to 44/220 on GPT 5.5 medium teacher rows and 46/220 on Qwen3.5-35B-A3B 8-bit rows.

The larger student runs are comparison points, not the answer to the 0.8B question. They show the same broad pattern: SFT teaches the loop, but it does not close the gap to the stronger teachers.

But the students did not inherit teacher-level SQL judgment; the stronger baselines remained far ahead. That is the honest result of this first post: protocol transfer, yes. Teacher-level decision making, no.

The natural next step is to keep the traces and the harness fixed but give the student a richer signal than one hard token per position. That is what the next post does: off-policy top-k soft-label distillation over these same teacher traces.

Your company. Your workspace.

Bring Nablo to your company.

Talk with us about the data your team works with and what you want Nablo to do.