Skip to content
Rota Nacional

Radar ·

TMAX trains terminal agents in Dockerized environments

A post published June 24, 2026 describes 14,600 Dockerized reinforcement-learning environments and an outcome-only DPPO recipe. It reports a 9B model scoring 27% on Terminal-Bench 2.0.

Published June 24, 2026, the post about TMAX describes 14,600 Dockerized reinforcement-learning (RL) environments and a DPPO recipe that uses outcomes alone to train terminal agents. It reports that a 9B model reached 27% on Terminal-Bench 2.0.

According to the post, the recipe and environment design may help engineers train agents for multi-turn shell tasks. To assess the result, consult the original publication and check how it defines the metric, benchmark, and test conditions; the summary figures alone do not establish comparability with other evaluations.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free