The paper describes a self-improvement method that evolves an agent alongside its evaluator. Evaluators stay fixed during each epoch; replacement happens only when a new evaluator performs better on reserved reference data (ground truth).\n\nThe proposal aims to reduce benchmark manipulation and preserve evaluation guarantees within each epoch. Organizations studying or applying this kind of method with AI should avoid sending personal or confidential data without appropriate authorization and controls.
Radar ·
Red Queen Gödel Machine coevolves agents and evaluators
A self-improvement method evolves an agent and its evaluator, keeping evaluators fixed during each epoch and replacing them only after stronger performance on held-out ground truth.
Rota Nacional
Bring privacy into your workflow.
30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.