Small-Model Distillation, Part 7: Training a 4B Model for Customer Support
TL;DR
Can a small language model get better at customer support by learning from successful conversations? I tried this with Qwen3.5-4B in the retail part of Tau2-bench, which simulates an online store. After one pass through the training data, the best trained version’s average success rate rose from 81.88% to 85.63% on 40 evaluation tasks.
Another version matched the training examples more closely but succeeded less often. Lower training loss did not identify the better agent. The measured improvement is still uncertain, and these evaluation scores also determined which version was selected. Performance on new tasks remains untested.
How Tau2-bench retail tasks work and are scored
Each retail task starts with a simulated customer asking for help, such as returning an item or changing a delivery address. The agent talks to the customer and uses tools to look up records and update orders. It must follow the store’s rules, identify the customer, and get confirmation before making changes that require approval.
The agent uses PydanticAI to connect the model to the store’s tools. At each turn, the model can reply to the customer or call a tool. Tau2-bench executes tool calls and returns their results, which inform the model’s next decision. The same model handles the support request throughout the conversation.
Tau2-bench checks each attempt in two ways:
- A database check verifies that the agent made the required changes to the records.
- A language judge checks whether the agent communicated the information required by the task.
An attempt passes only when both checks pass. For example, an agent can tell the customer it exchanged the correct item but send the wrong item identifier to the tool. The database check catches that mistake even if the reply sounds correct.

DeepSeek V4 Flash served three roles through separate calls: generating teacher examples, simulating customers, and judging replies. During student evaluation, Qwen handled the support request while DeepSeek supplied the customer and judge.
Conversations that passed both checks could supply training examples. The student learned the support agent’s responses and tool calls, with customer messages and tool results as context.
Training data from demonstrations and recoveries
The 114 retail tasks were split into 74 for training and 40 for evaluation.
DeepSeek and the original Qwen model each attempted the 74 training tasks twice. Conversations that passed both the database check and the language judge supplied examples of successful behavior. The training data therefore combined teacher demonstrations with Qwen’s own correct responses and tool calls.
For some failed Qwen attempts, DeepSeek took over with the conversation so far and the store records exactly as Qwen had left them. These continuations supplied examples of how to finish a request after a mistake, provided the completed attempt passed both checks.
The final dataset contained 2,357 examples from 72 of the 74 training tasks. Each example paired the conversation so far with one agent response or tool call to learn from. Of these, 1,379 came from DeepSeek demonstrations, 926 from Qwen’s own successful conversations, and 52 from DeepSeek taking over failed attempts.
The other two tasks contributed no examples because none of their attempts passed both checks. During DeepSeek’s recovery attempts, one still failed the database check; the other passed that check but failed the language judge.
One conversation becomes a training example
This shortened teacher conversation shows how an example is formed. The simulated customer has supplied a name and ZIP code, and a tool has returned the matching account identifier.
| Context the model receives | From the saved conversation |
|---|---|
| Customer message | “Sure! My name is Yusuf Rossi, and my zip code is 19122.” |
| Previous tool call | find_user_id_by_name_zipName: Yusuf Rossi; ZIP: 19122 |
| Tool result | yusuf_rossi_9620 |
The training target is the next agent turn, which includes this reply:
Thank you, Yusuf! I found your account. Let me pull up your details to verify your order.
That same turn calls:
get_user_details(user_id="yusuf_rossi_9620")
The example teaches a small step: use the account identifier returned by one tool in the next tool call. The input also includes the earlier conversation and store rules, omitted here for brevity. Only the next agent turn is a target; the preceding customer message and tool result provide context.
Training Qwen with supervised fine-tuning
Supervised fine-tuning (SFT) trained Qwen to produce the saved agent response or tool call for each example. The examples were collected before training and stayed fixed throughout the run.
This method is called hard-token trajectory SFT. A trajectory is the sequence of messages and tool calls in an attempt; a token is a piece of text the model reads or produces. Hard-token training learns from saved text and tool calls rather than the teacher’s probabilities for possible next tokens. Here, distillation uses the teacher’s responses alongside the smaller model’s own successful responses.
The loss measured how well Qwen reproduced the next agent turn given the store’s rules, available tools, and conversation so far. Earlier messages and tool results did not contribute to that loss.
Both training runs used LoRA, which trains a small set of additional weights while keeping the original model weights fixed. The two versions used the same data and differed only in learning rate, which controls the size of the training updates: 0.00005 for Candidate A and 0.0001 for Candidate B.
| Setting | Value |
|---|---|
| Starting model | Qwen3.5-4B |
| Passes through the dataset | 1 |
| Training steps | 295 |
| Examples per training update | 8 |
| LoRA settings | Rank 16, alpha 16, no dropout |
| Numerical precision | BF16, without quantization |
| Optimizer | AdamW, weight decay 0.01 |
| Learning-rate schedule | Warmup over the first 5% of steps, followed by cosine decay |
| Random seed | 3407 |
Each run took about four hours.
Evaluation results
Each model attempted the same 40 evaluation tasks four times. Repeating tasks helps show whether the model succeeds consistently, since its actions and the simulated customer’s replies can change between attempts.
- Pass^1: the average success rate across attempts.
- Pass^4: the percentage of tasks where all four attempts succeeded.
| Model | Pass^1 | Pass^4 |
|---|---|---|
| DeepSeek teacher | 92.50% | 82.50% |
| Original Qwen | 81.88% | 62.50% |
| Candidate A, learning rate 0.00005 | 85.63% | 70.00% |
| Candidate B, learning rate 0.0001 | 80.00% | 65.00% |
Candidate A had the highest success rate of the Qwen versions, improving by 3.75 percentage points over the original model. Across the 40 tasks, it succeeded more often on nine, less often on six, and equally often on 25.
Candidate B had lower training loss, meaning it matched the training examples more closely. Its evaluation success rate nevertheless fell to 80.00%, below the original Qwen model. Choosing between the trained versions by training loss alone would have selected Candidate B.
Candidate A’s estimated gain of 3.75 percentage points had a 95% confidence interval from −3.75 to +11.25 points across all 40 tasks. A paired bootstrap estimated this interval by repeatedly sampling tasks from the evaluation set, keeping each task’s results for both models together. The range includes no improvement, so this sample does not give a clear answer about whether the gain would hold on other tasks.
A secondary comparison used 39 tasks, leaving out one that required changing an order’s payment method. That operation was absent from the training tasks’ reference solutions. Without it, the gain was 3.85 points, with an interval from −3.85 to +11.54 points. The conclusion stayed the same. An operation missing from training is still a valid challenge, so the main result includes all 40 tasks.
I selected Candidate A using these evaluation scores, so the 40 tasks did not provide an independent final test.
What changed in two conversations
In one evaluation task, the customer wanted to exchange a pair of earbuds for the cheapest other earbuds already in their order. The original Qwen chose a cheaper pair from the product catalog, but that pair was not in the order. Candidate A chose the correct pair from the order. Both conversations passed the language judge, but only A passed the database check.
| Model | Item identifier sent to the exchange tool |
|---|---|
| Original Qwen | 8555936349, a different catalog item |
| Candidate A | 1646531091, the requested item already in the order |
A refund task illustrates a regression. The customer wanted to return three items and receive the refund through PayPal, although they had paid by Visa. The original model explained the restriction and transferred the request to a person when the customer insisted.
Candidate A tried to send the refund to PayPal. The tool returned:
Error: Payment method should be the original payment method
Candidate A then processed the return to the card without asking the customer again. The customer objected, and the attempt failed the database check even though it passed the language judge. On this task, the original model passed all four attempts and A passed three.
GPU time and API costs
The experiment incurred two kinds of cost: GPU rental to train and run Qwen, and DeepSeek API charges for the teacher, simulated customer, and language judge.
Training Candidate A took about 3.8 hours and Candidate B about 3.9 hours, for a total of 7.7 GPU hours. Evaluating the original Qwen model took another 2.3 hours for 160 attempts. Loading and evaluating both trained versions took about 2.6 hours for their combined 320 attempts.
The recorded times do not provide a complete GPU rental bill. Multiplying training time by an hourly rate would omit setup and evaluation time, so the available records cannot establish the experiment’s total dollar cost.
The DeepSeek API charges were recorded separately. The amounts below are rounded and cover the evaluation runs:
| Evaluation | Recorded API cost |
|---|---|
| DeepSeek teacher, 40 tasks × 4 attempts | $0.34 |
| Original Qwen, 40 tasks × 4 attempts | $0.074 |
| Candidate A, 40 tasks × 4 attempts | $0.071 |
| Candidate B, 40 tasks × 4 attempts | $0.072 |
For the teacher runs, these charges include DeepSeek acting as the support agent, customer, and judge. For the Qwen runs, they cover only the DeepSeek customer and judge. Qwen ran on the rented GPU, whose cost is not included in the table. The table also excludes API calls used to collect the training examples.
Because the benchmark pays a model to simulate the customer, these evaluation charges do not directly represent the cost of serving real customers.
Lower training loss did not identify the better agent
Candidate A achieved a higher average success rate in this evaluation, although the uncertainty interval includes no improvement. Candidate B fit the saved training examples more closely and still performed worse than the original model. The full conversations and checks of the resulting store records revealed a difference that training loss could not. Whether Candidate A’s gain carries over to new tasks remains unanswered.