Staff SRE, AI Infrastructure
Core
Build and operate large-scale GPU compute, Kubernetes, and cloud infrastructure platforms to support Wayve's AI development for autonomous driving.
Role type
Staff Site Reliability Engineer (AI Infrastructure) (via careerplan.io/jobs/085402fd-1976-4725-bfc3-f0e9dd3bc2ab-staff-sre-ai-infrastructure-at-wayve)
Builds
Reliable access to large-scale GPU compute, Kubernetes, storage, and deployment systems for AI training and inference.
Domain
Autonomous driving / AI Infrastructure / Cloud Computing
Deliverable
infrastructure
Required skills
Large-scale cloud infrastructure ownership, Kubernetes administration, Python or Go programming, Infrastructure as Code, CI/CD and GitOps, Distributed systems failure analysis, GPU/ML-training infrastructure experience, Capacity planning and forecasting, Observability and SLO management.
Preferred skills
Experience with Argo CD or Flux, Experience with HPC infrastructure, Experience with vehicle workshops and labs.
Responsibilities
Identify reliability risks and engineer durable solutions for recurring operational problems, Improve observability, automation, deployment safety, and incident learning processes, Troubleshoot complex failures spanning compute, networking, storage, and distributed workloads, Develop continuous GitOps delivery pipelines, Enhance Kubernetes scheduling, autoscaling, and resource efficiency, Build forecasting models for GPU and storage capacity.