Dom Colligan
I like evaluations, data and post-training.
Economics, Finance and Data Science at Imperial College London.
Work
An RL environment which trains a model to write test suites for functions which operate on structured data. Post-trained Qwen3.5-4B on it with GRPO: 20 steps took it from 0.41 to 0.88 on the held-out split, level with an untrained 8B on the families it trained on. I evaluated the checkpoint twice and the two runs disagree by 0.13 on the unseen family, which is written up rather than tidied away.
A pre-registered study of when telling an agent the rules substitutes for enforcing them. Eight models, four open-weight families, a stronger and a weaker sibling in each. Telling raises the first-action safe rate in every family, but enforcement still adds on top of telling, and it is nearly all of the barrier in the weak families.
39 tasks over an agent harness, with a deterministic offline grader and every trajectory released. Built to separate two things that usually get measured together: what the model was told, and what the runtime would actually let it do.
Core ML implemented in pure Python from first principles, working up from gradient descent toward autograd and networks.
Now
July 2026. AI and Data Science Intern at Let Tech Solutions, where I own an evaluation for a production agent workflow. Research assistant at the IDEA Lab, Imperial.