Four reward signals for programming tasks can fail to reflect human intent. Learn how to use them as starting points for more robust evaluations.
The material describes four reward signals for programming tasks: test suites, scored checklists for web pages, real conversations between engineers and assistants, and project evaluation by an agent.
Each signal offers a way to assess a coding agent’s work, but may fail to capture what people actually intended. This gap creates room for reward hacking: optimizing the score without meeting the expected goal.
When studying these examples, compare what each signal measures with the task’s intent. Use them to spot potential gaps and plan additional checks; do not assume that any single signal proves the result is good.
If you use AI to analyze conversations, code, or internal projects, follow your organization’s data policy. Avoid sending personal or confidential information without authorization, and review outputs before using them in an evaluation.