Published on June 29, 2026, DiscoBench is described as a benchmark that tests LLM search agents on vague, underspecified, or factually incorrect requests. The description provides no performance results.
Published on June 29, 2026, DiscoBench is a benchmark for search agents powered by language models. It evaluates whether agents can handle vague, underspecified, or factually incorrect requests during information retrieval and multi-step reasoning. According to the description, the benchmark is intended to help engineers examine agent behavior on real-world search tasks.
The description provides no scores, comparative results, or protocol details. To verify the scope and any results, consult the original publication and check its task definitions, evaluation criteria, and reported data; performance conclusions cannot be inferred from this summary alone. If using AI to study or apply the benchmark, avoid submitting sensitive organizational data without authorization.