Coding-agent harnesses improve over ten iterations
A paper describes a framework using observable components, condensed trajectory evidence, and decisions tested against task results. On Terminal-Bench 2, pass@1 rises from 69.7% to 77.0% after ten iterations.
Published on April 29, 2026, the paper presents a framework for automatically evolving coding-agent harnesses. Its approach combines observable components, condensed evidence from trajectories, and decisions tested against task outcomes; the aim is to make harness changes easier to evaluate, attribute, and revert.
On Terminal-Bench 2, the reported pass@1 rises from 69.7% to 77.0% over ten iterations. To verify the result, consult the original paper and check how it defines the benchmark, iterations, and pass@1 metric, as well as the results it reports.