Software Engineer, Reliability
Core
Design and implement solutions to ensure the scalability, stability, and performance of OpenAI's rapidly evolving infrastructure for millions of users.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Resilient, scalable infrastructure supporting AI research and deployment products
Domain
Cloud infrastructure & AI systems
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure, container orchestration (Kubernetes), Infrastructure as Code (Terraform/CloudFormation), observability (DataDog, Prometheus, Grafana, Splunk), microservices architecture, service mesh, fault-tolerant design patterns, SLO/SLI management, automation tooling
Preferred skills
Chaos engineering, load testing, synthetic testing, CPU/storage/GPU lifecycle management
Technologies
Kubernetes, Terraform, CloudFormation, DataDog, Prometheus, Grafana, Splunk
Responsibilities
Design scalable infrastructure solutions, build load/chaos/synthetic testing software, maintain automation tools for repetitive tasks, manage CPU/storage/GPU/network lifecycle platforms, implement fault-tolerant design patterns, develop and maintain SLOs/SLIs, participate in on-call rotation for critical incidents
Seniority
Senior, hands-on IC