CareerPlanSign in

Site Reliability Engineer

🌐 Remote💼 Full-time💰 $150,000–$150,000🗓 2026-07-10 → 2026-09-25

Core

Ensure the stability, resilience, and operational excellence of Runpod's global AI developer cloud platform serving millions of developers.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Reliability frameworks, observability systems, automation tools, and production hardening for distributed AI infrastructure.

Domain

AI Infrastructure / Cloud Computing / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Linux systems administration, Networking, Container orchestration, Distributed systems architecture, SLI/SLO definition, Incident response leadership, Python/Go/Bash scripting, Monitoring and alerting design

Preferred skills

GPU infrastructure management, AI/ML platform experience, High-scale environment reliability, Infrastructure as Code, Startup environment experience, Internal reliability platform development

Technologies

Prometheus, Grafana, Python, Go, Bash

Responsibilities

Define and implement SLIs/SLOs for critical services; Lead incident response and coordinate cross-team mitigation; Conduct blameless postmortems; Design and improve monitoring and alerting systems; Automate recurring operational workflows; Partner with engineering teams to improve system resilience.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.