Site Reliability Engineer
Core
Ensure the stability, resilience, and operational excellence of Runpod's global AI developer cloud platform serving millions of developers.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Reliability frameworks, observability systems, automation tools, and production hardening for distributed AI infrastructure.
Domain
AI Infrastructure / Cloud Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Networking, Container orchestration, Distributed systems architecture, SLI/SLO definition, Incident response leadership, Python/Go/Bash scripting, Monitoring and alerting design
Preferred skills
GPU infrastructure management, AI/ML platform experience, High-scale environment reliability, Infrastructure as Code, Startup environment experience, Internal reliability platform development
Technologies
Prometheus, Grafana, Python, Go, Bash
Responsibilities
Define and implement SLIs/SLOs for critical services; Lead incident response and coordinate cross-team mitigation; Conduct blameless postmortems; Design and improve monitoring and alerting systems; Automate recurring operational workflows; Partner with engineering teams to improve system resilience.
Seniority
Senior, hands-on IC
