<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Dom Colligan</title>
  <subtitle>Write-ups on RL environments, graders and evals.</subtitle>
  <link href="https://dtcolligan.com/feeds/atom.xml" rel="self"/>
  <link href="https://dtcolligan.com/"/>
  <id>https://dtcolligan.com/</id>
  <updated>2026-09-04T00:00:00Z</updated>
  <author><name>Dom Colligan</name></author>
  <entry>
    <title>Suitesmith: training a 4B model to write test suites</title>
    <link href="https://dtcolligan.com/posts/suitesmith/"/>
    <id>https://dtcolligan.com/posts/suitesmith/</id>
    <updated>2026-09-04T00:00:00Z</updated>
    <summary>An RL environment which trains a model to write test suites for functions which operate on structured data. 20 steps of GRPO took a 4B from 0.41 to 0.88, and evaluating the same checkpoint twice disagreed by 0.13 on the tier that mattered.</summary>
  </entry>
</feed>
