Staff Cloud SRE – AI/ML Platform & GPU Compute
Core
Founding Staff SRE building reliability foundations for large-scale AI cloud platforms, including Model Development and GPU Compute fleets.
Role type
Staff Cloud Site Reliability Engineer (AI/ML Infrastructure)
Builds
Model Development Platform and multi-tenant GPU compute fleets for model training and inference
Domain
Artificial Intelligence / Machine Learning / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
SRE/Production Engineer experience, GPU-backed environment operations, MLOps pipeline management, Kubernetes production operations, AWS/GCP/Azure production workloads, distributed systems management, Linux fundamentals, Python/Go/C++ scripting, deep troubleshooting, observability stack design
Preferred skills
Infrastructure-as-code (Terraform), SLO/SLI definition, founding SRE experience, leadership in growing SRE functions
Technologies
Kubernetes, AWS, GCP, Azure, Python, Go, C++, Datadog, Prometheus, Grafana, OpenTelemetry
Responsibilities
Own reliability/availability/performance of Model Dev and GPU platforms, define and operationalize SLOs/SLIs, lead incident response and root cause analysis, design observability systems, build automation for cluster operations and CI/CD
Seniority
Staff, hands-on IC with founding responsibilities