A report on reinforcement learning says that focusing on beneficial behaviors improved alignment across domains and resisted adversarial pressure.
A report from OpenAI says that reinforcement learning focused on beneficial behaviors in realistic scenarios improved alignment across domains and resisted adversarial pressure. It also says the training combined a small share of behavior-focused data with other data, without detailing the full composition here.
The results suggest this kind of training may generalize beyond the domains represented in the training data. When using AI to study or apply the research, avoid submitting personal data or internal organizational information without authorization, and follow applicable data-protection rules.