Research

Notes from the loop.

What we learn building autonomous systems for RL environments: reward hacking, verifiers, data and the failures along the way.

No posts yet.The first ones are on the way.