Research Focus
I’m a 20-year-old self-taught independent researcher working on AI. I’m interested in reinforcement learning post-training methods for language models, especially those which make use of formal reasoning and executable verification to improve abilities at math and code. I left college to focus on independent research.
Recently I developed and released a C verification environment that uses Frama-C to check C programs against their formal specifications, and a bounded integer math environment which checks answers in Lean, where the model supplies an integer answer rather than a proof. I also worked on grammar-constrained generation of proof tactics in Lean. I typically release code, task suites, and experimental evidence for the tasks I work on.
Research Experience
- Build RL environments for mathematical reasoning and code generation where rewards come from executable checks and formal verification.
- Build isolated, fail-closed judges and test them against malformed outputs, timeouts, crashes, inconsistent verifier reports, and attempts to exploit the reward.
Selected Projects & Technical Writing
- Built an RL environment where models write C functions and reward comes from Frama-C proving that the implementation satisfies a fixed specification, including runtime safety.
- Released 64 tasks, an isolated Frama-C judge, adversarial tests, and reproducible proof records. The published source passed 57/57 tests.
- Article: An RL Environment Where C Code Has to Be Proved, Not Just Tested.
- Built a math RL environment for bounded integer problems that does not store expected answers. Models submit integers or complete finite certificates, and reward comes from checking each submission against the full encoded problem rather than requiring a model-written proof.
- Article: Grading Mathematical Answers Without Answer Keys.