HERMES: component-level harness with LLMs for software engineering tasks
An arXiv paper proposes HERMES, a code-agent harness in which each repository component has a resident LLM. The authors report gains in migration and on four software engineering benchmarks against comparable baseline harnesses.
An arXiv paper proposes HERMES, a code-agent harness in which each repository component has a resident LLM that knows its code. According to the published summary, the authors report gains in migration tasks and on four software engineering benchmarks, compared with comparable baseline harnesses. The central claim is that harness design, not only model choice, can produce large gains in long-horizon code agents. The source does not give benchmark names, figures or implementation details, which should be checked in the original paper.
Relevance for Rota Nacional users: Rota does not offer HERMES or reproduce the architecture described; the material is a research reference. Anyone using AI to study the paper or test similar ideas should avoid pasting proprietary code, credentials or personal data into prompts. Rota detects CPF, CNPJ, e-mail, phone and person names before any model runs and applies the organization's policy (placeholders, removal or block), but this does not replace reviewing secrets and sensitive code before sending. To verify the claims, read the original paper and reproduce only the experiments you are authorized to run.