Skip to content
Rota Nacional

Radar ·

ROME explores segment-level credit for agent training

A 3-billion-parameter model uses segment-level rewards, rather than assigning credit only by token or trajectory, to train agents that use tools.

The publication describes ROME, a 3-billion-parameter model trained with IPA and rewards assigned by segment. Its process includes pretraining on structured coding tasks, supervised fine-tuning with error masking, and reinforcement-learning optimization at the segment level.

The approach offers an alternative to token-level or trajectory-level credit assignment for training tool-using agents. The material gives no comparative results; organizations studying or applying the approach with AI should avoid submitting code, credentials, or identifiable internal data, and use only authorized content.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free