TMAX trains terminal agents in Dockerized environments
A post published June 24, 2026 describes 14,600 Dockerized reinforcement-learning environments and an outcome-only DPPO recipe. It reports a 9B model scoring 27% on Terminal-Bench 2.0.
Published June 24, 2026, the post about TMAX describes 14,600 Dockerized reinforcement-learning (RL) environments and a DPPO recipe that uses outcomes alone to train terminal agents. It reports that a 9B model reached 27% on Terminal-Bench 2.0.
According to the post, the recipe and environment design may help engineers train agents for multi-turn shell tasks. To assess the result, consult the original publication and check how it defines the metric, benchmark, and test conditions; the summary figures alone do not establish comparability with other evaluations.