A paper examines agents that analyze failures and propose small harness changes, keeping them only after regression tests pass. It reports higher success on unseen Terminal-Bench-2.0 data.
Published on June 10, 2026, the Self-Harness paper explores agents that analyze their own failures and suggest small changes to a harness. Changes are kept when they pass regression tests. The approach aims to improve prompts, tools, retries, and verification without fine-tuning.
The summary reports improved success rates on unseen Terminal-Bench-2.0 data with MiniMax, Qwen, and GLM. To assess the result, consult the original paper and verify its method, regression tests, and experimental conditions; the available summary provides no metrics or further details.