Senior Software Engineer, RL Post-Training Frameworks
This role involves designing and building scalable reinforcement learning (RL) post-training infrastructure that supports the full lifecycle of training-inference-rollout loops across heterogeneous hardware. The engineer will contribute to open-source RL frameworks, optimize distributed systems for performance and fault tolerance, and collaborate with AI researchers, infrastructure teams, and hardware partners to advance large-scale AI training. Work includes improving runtimes like Ray and Monarch, integrating high-performance inference engines, and enabling efficient coordination of actor, critic, and reward models at scale.