Published on September 17, 2026, the report describes RULER in OpenPipe ART: an LLM ranks trajectories using plain-English criteria, and the comparisons are converted into rewards for GRPO training.
In a post published on September 17, 2026, the author presents RULER in OpenPipe ART. Engineers define evaluation criteria in plain English; an LLM ranks agent trajectories, and ART converts the comparisons into rewards used for GRPO training. The approach aims to evaluate multi-step behavior without manually writing a set of scoring rules. The report says it was used with a Qwen3 1.4B agent playing 2048.
To check the finding, consult the project’s original publication and verify how the criteria and trajectories were defined, as well as how the comparisons became rewards. The available text gives no comparative metrics or quantitative results, so it does not by itself establish that the method outperforms other approaches.