Salesforce presents UserRL, a framework that uses standardized gym environments and simulated users to train agentic models. The account discusses SFT cold starts, trajectory rewards, and scalable training with simulated users.
When studying the work, separate these three training choices and note which matches the question you want to investigate: initialization, trajectory evaluation, or training scale.
To apply similar ideas, first define the multi-turn interaction you want to study and what observations would allow you to compare results. Treat these choices as avenues for investigation, not proof that the framework solves any particular task.
If you use AI to study or adapt the material, avoid entering personal, confidential, or internal data without authorization. Prefer synthetic or properly de-identified examples, and check conclusions against the original source.