Prompt Optimization on Tasks Nablo Generated Itself
TL;DR
Nablo’s workspace agent answers questions over a customer’s SQLite database. For this experiment we let Nablo generate its own evaluation tasks from the database, then ran prompt optimization against those tasks. On nine tasks the optimizer never saw, the pass rate went from 2 of 9 to 6 of 9. The winning change was an eight-line edit to the agent’s instructions, with no database-specific details in it. Everything ran locally on one machine with one open-weight model.
The agent and its job
The workspace agent connects to a customer database — here, a demo customer’s SQLite banking database. Given a question, it can inspect the schema, run SQL, read the results, and submit an answer. It has a fixed budget of turns. A task counts as solved only when the submitted answer matches a verified reference result that the agent never sees.
With its baseline instructions, the agent solved 2 of 9 unseen tasks. It knew how to use its tools, but its strategy for reaching a verifiable answer was weak.
Evaluation data, generated automatically
Nablo’s task generation pipeline works from the database alone. It drafts candidate questions, executes each reference query against the real database, and keeps only tasks whose reference answer verifies. Near-duplicates are removed, and any task whose reference result is too large for the agent to read is rejected — a task the agent cannot actually see through is not a fair test. Thirty-eight tasks survived this filter, split three ways: 19 to optimize on, 10 to validate candidate instructions, and 9 held out untouched until the very end.
This is the point of generated evaluation: it measures real behavior on this customer’s actual database, not performance on a generic benchmark.
Optimizing the instructions, carefully
Prompt optimization means improving the written instructions an agent starts with, rather than its weights. We used GEPA, an open-source method that works in a loop: run the agent on training tasks, collect the failures together with feedback on why they failed, then ask a language model to propose the smallest possible edit to the instructions. Edits that raise the validation score are kept; the rest are discarded.
Our first run found a loophole. The top instructions scored 0.6 on validation, but they had hardcoded table names from the database. That is copying, not skill. The score was real, but the instructions would collapse on any new task set. The fix was not more data but rules: every proposal now carries constraints that forbid copying table or column names and example details, demand rules that generalize, and require the smallest possible edit. The untouched holdout is graded only once, at the end, for both the baseline and the winner.
Results
With those guardrails in place, the rerun looked like this:
| Instructions | Validation (10 tasks) | Holdout (9 tasks) |
|---|---|---|
| Baseline | 3/10 | 2/9 (22%) |
| Optimized | 4/10 | 6/9 (67%) |
The winning instructions are eight added lines about how to work: planning before querying, checking edge cases in aggregations, and reading back results before answering. They contain zero table or column names from the database, so they transfer rather than memorize.
The same local model played every role. One Qwen3.6-35B model, 4-bit, where each weight is compressed to four bits, served from the Mac, drafted the tasks, solved them, and proposed the instruction edits. The model improved its own operating instructions. Scoring stayed deterministic — executed SQL compared against verified references, never a model’s opinion.
Honest limits
Nine holdout tasks is a small sample. Six of nine versus two of nine is a real gap, but each task is worth 11 points, so treat 67% as a reading with wide error bars, not a precise number. Roughly a quarter of attempts end uncertain rather than pass or fail — the reference answer and the agent’s answer could not be confidently compared — and that verification ceiling is the next thing to attack.
What this means
The loop is closed and fully local: the customer’s database produces a verified benchmark, the optimizer adapts the agent to it, and the new instructions are adopted only when the score on unseen tasks improves. No model was retrained, no data left the machine, and nothing in the winning instructions is specific to this customer. Next steps are larger holdouts and sharper verification of free-form answers.