Published on April 25, 2026, PARE-Bench contains 143 tasks for testing agents in stateful app simulations, including goal inference, intervention timing, and orchestration across apps.
Published on April 25, 2026, PARE-Bench evaluates proactive agents in stateful app simulations. The PARE method represents apps as finite-state machines to simulate active users. The benchmark includes 143 tasks across communication, productivity, calendar, and lifestyle apps.
The tasks test three capabilities: inferring goals, deciding when to intervene, and orchestrating actions across apps. The work gives engineers a benchmark for evaluating agents in sequential, stateful interactions. To verify the details, consult the original publication and check its task definitions, simulated environments, and evaluation criteria; the available summary does not specify quantitative performance results.