Microsoftās coding-agent skill distillation beats GEPA at prompt optimization
A new Microsoft paper says a coding agent reading saved logs can write better system prompts than leading tuning tools for about $1.60 a pass
In a paper titled “Coding Agents are Strong Prompt Optimizers,” published on September 23, 2026, Microsoft researchers introduce Coding-Agent Skill Distillation, or CASD. The method writes improved system prompts directly from a static pile of saved agent logs, and it does so for approximately $1.60 per pass.
That price matters because the competition is not cheap. According to the paper, CASD comes in over 22 times cheaper than validation-gated iterative methods, the trial-and-error approach that has become standard for tuning agent prompts.
How CASD works, and how it scored
CASD needs no iterative search, no live interaction with the environment, and no held-out validation set.
Instead, a coding agent takes in an entire corpus of saved trajectories, which are records of what an agent did on past tasks. It runs a single corpus-wide statistical analysis, pulling out aggregate statistics, recurring failure modes, and representative episodes. From that analysis, it distills behavioral rules and writes them into an optimized prompt.
The researchers frame this as reading the whole corpus rather than small minibatches. Iterative optimizers typically look at slices of data at a time, which can make it harder to spot patterns that only show up across many episodes.
The headline result is an average gain of +16.6 percentage points over the unoptimized baseline. CASD also beat validation-gated reflective search in the style of SkillOpt across its metrics.
Against GEPA, a widely used reflective prompt optimizer, CASD won on three of four agentic benchmarks:
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
- ALFWorld: CASD 83.3 vs. GEPA 74.0
- ϲ-bench retail: CASD 40.0 vs. GEPA 39.2
- ϲ-bench telecom: CASD 39.2 vs. GEPA 17.5
- SpreadsheetBench-Verified: CASD 51.3 vs. GEPA 60.7
The paper also says CASD kept its lead over rival methods even when those rivals had extra validation data and unrestricted environment access.
Why GEPA is the benchmark that matters
The research describes GEPA as widely adopted at companies including Databricks, Shopify, Dropbox, and OpenAI, and credits it with substantial gains in open-source work.
The paper comes from Microsoft researchers Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, and Sumit Gulwani. It builds on related Microsoft work called SkillOpt, which treats agent skills as trainable text parameters and was introduced in a paper published in June 2026.
SkillOpt represents the validation-gated, iterative school of thought. CASD is effectively Microsoft arguing that, in many cases, its own earlier approach is more work than necessary.
What this means for teams building agents
If a single analysis pass at approximately $1.60 can match or beat iterative optimization on most tested tasks, the case for expensive tuning loops gets harder to make.
Dropping the need for held-out validation data removes a significant preparation burden. Skipping live environment interaction lowers risk too, since iterative optimizers that test prompts against real systems can trigger real actions when the agent handles customer accounts or production tools.
The SpreadsheetBench-Verified result, where GEPA finished 9.4 points ahead, shows a reflective, iterative optimizer can still win on certain structured tasks. The method also depends on the quality of the saved trajectories, so a corpus full of narrow or unrepresentative episodes could produce rules that do not generalize.