Skip to content
Rota Nacional

Guides ·

UserRL trains agents with simulated users

Salesforce presents UserRL, a framework that combines standardized gym environments and simulated users to train agentic models. The account highlights choices involving SFT cold starts, trajectory rewards, and scalable training.

Salesforce presents UserRL, a framework that uses standardized gym environments and simulated users to train agentic models. The account discusses SFT cold starts, trajectory rewards, and scalable training with simulated users.

When studying the work, separate these three training choices and note which matches the question you want to investigate: initialization, trajectory evaluation, or training scale.

To apply similar ideas, first define the multi-turn interaction you want to study and what observations would allow you to compare results. Treat these choices as avenues for investigation, not proof that the framework solves any particular task.

If you use AI to study or adapt the material, avoid entering personal, confidential, or internal data without authorization. Prefer synthetic or properly de-identified examples, and check conclusions against the original source.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free