Posts
-
September 4, 2026
An RL environment which trains a model to write test suites for functions which operate on structured data. 20 steps of GRPO took a 4B from 0.41 to 0.88, and evaluating the same checkpoint twice disagreed by 0.13 on the tier that mattered.