MemoHarness adapts harnesses with prior executions
Published on July 17, 2026, MemoHarness adapts six dimensions of agent harnesses and retrieves lessons from similar past cases. On a shell-agent benchmark, it scored 0.806, versus 0.722 for the strongest fixed-harness baseline.
MemoHarness modifies six dimensions of an agent harness and retrieves lessons from previous executions similar to a new case. The proposal explores improving agent control layers using their own executions, without labels at test time or additional search.
In the shell-agent benchmark reported by the source, the method reached 0.806, while the strongest fixed-harness baseline scored 0.722. To check the result and understand the scope of the comparison, consult the original MemoHarness material and verify how the benchmark defines its metric, evaluated cases, and baseline.