Senior Cloud SRE - AI/ML Platform & GPU Compute
Core
Build and scale the reliability foundations of an AI cloud platform, including Model Development and GPU Compute fleets for automated driving systems.
Role type
Founding Senior Cloud Site Reliability Engineer (AI/ML Infrastructure)
Builds
Model Development Platform and large-scale, multi-tenant GPU compute clusters for model training and inference.
Domain
Autonomous driving / Embodied AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes production operations, AWS/GCP/Azure cloud platforms, distributed systems troubleshooting, Linux fundamentals, Python/Go/C++ scripting, observability stack design (Prometheus/Grafana/OpenTelemetry), incident response and root cause analysis, capacity planning for large clusters.
Preferred skills
GPU-backed environment operations, MLOps pipeline experience, Infrastructure-as-Code (Terraform), SLO/SLI definition and reliability program building, founding SRE experience.
Technologies
Kubernetes, AWS, GCP, Azure, Python, Go, C++, Prometheus, Grafana, OpenTelemetry, Terraform
Responsibilities
Own reliability, availability, and performance of Model Dev Platform and GPU Compute; define and operationalize SLOs/SLIs; lead incident triage and post-mortems; design monitoring, logging, and tracing systems; build automation for cluster operations and self-healing patterns; improve CI/CD safety and deployment velocity.
Seniority
Senior, hands-on IC (Founding role)