Staff Cloud SRE – AI/ML Platform & GPU Compute
This role involves building and scaling the reliability foundations of Wayve's AI cloud platform, focusing on the Model Development and GPU Compute environments. The engineer will define SRE frameworks, automation, and operational standards while ensuring resilient, efficient infrastructure for large-scale AI training and inference. The position bridges AI research, cloud infrastructure, and production operations, with a strong emphasis on observability, incident response, and platform ownership.