Staff ML Performance Engineer (Inference Optimisation)
Optimise machine learning inference performance for edge accelerators and GPUs, focusing on efficient execution of large transformer models on low-power devices. Profile bottlenecks across the full stack—from model graph to kernel execution—and implement compiler, runtime, and kernel-level optimisations. Build benchmarking systems and collaborate with model teams to influence architecture for on-device efficiency.