Senior Software Engineer, RL Post-Training Frameworks
This role involves designing and building scalable reinforcement learning (RL) post-training infrastructure that operates efficiently from single-GPU experiments to large-scale production systems. You'll optimize distributed training-inference-rollout loops across GPUs, CPUs, and LPUs, contribute to open-source RL frameworks like VeRL and TorchTitan, and collaborate with hardware and research teams to enhance system performance and reliability. The position emphasizes solving complex distributed systems challenges in AI, including fault tolerance, elastic scaling, and integration with next-generation hardware.