Published on June 24, 2026, the briefing describes a benchmark based on 90 tasks from papers in the Nature family. According to the post, the top agent exceeded the published reference result on 17.8% of tasks.
NatureBench evaluates coding agents on 90 tasks based on papers in the Nature family. It compares their results with published findings, without web search or access to the original method. According to the post, the top agent exceeded the published SOTA on 17.8% of tasks. The benchmark aims to assess whether agents can select methods and reproduce research results.
To consult and verify the finding, locate the original post and check how it defines the tasks, comparison, and metric; the briefing does not provide those details. If you use AI to study or apply the methods, avoid entering internal or identifiable organizational data, or follow your team’s data policy before submitting it.