Argus reports 78% on SWE-Bench Pro with a verified runtime
Published on August 8, 2026, the account describes Argus, a runtime for long-horizon agentic tasks, and reports 78% on SWE-Bench Pro versus 59% for Direct Copilot.
Published on August 8, 2026, the post presents Argus as a general-purpose agentic reasoning runtime for long-horizon tasks with loosely defined goals and sparse feedback. It says the system conditions pivots on verification and reports 78% on SWE-Bench Pro, compared with 59% for Direct Copilot.
The post suggests that engineers building long-running agents evaluate how verification controls changes to runtime state. To interpret the comparison, consult the original post and check its methodology, evaluation conditions, and benchmark scope; those details are not included in the available summary.