ROME explores segment-level credit for agent training
A 3-billion-parameter model uses segment-level rewards, rather than assigning credit only by token or trajectory, to train agents that use tools.
The publication describes ROME, a 3-billion-parameter model trained with IPA and rewards assigned by segment. Its process includes pretraining on structured coding tasks, supervised fine-tuning with error masking, and reinforcement-learning optimization at the segment level.
The approach offers an alternative to token-level or trajectory-level credit assignment for training tool-using agents. The material gives no comparative results; organizations studying or applying the approach with AI should avoid submitting code, credentials, or identifiable internal data, and use only authorized content.