CareerPlanGet AI match score →

Software Engineer, Reliability

San Francisco💼 Full-time🗓 2025-10-17 → 2026-07-31

Core

Design and implement solutions to ensure the scalability, stability, and performance of OpenAI's rapidly evolving infrastructure for millions of users.

Role type

Senior IC Site Reliability Engineer (SRE)

Builds

Resilient, scalable infrastructure supporting AI research and deployment products

Domain

Cloud infrastructure & AI systems

Deliverable

production ML models | infrastructure

Required skills

Cloud infrastructure, container orchestration (Kubernetes), Infrastructure as Code (Terraform/CloudFormation), observability (DataDog, Prometheus, Grafana, Splunk), microservices architecture, service mesh, fault-tolerant design patterns, SLO/SLI management, automation tooling

Preferred skills

Chaos engineering, load testing, synthetic testing, CPU/storage/GPU lifecycle management

Technologies

Kubernetes, Terraform, CloudFormation, DataDog, Prometheus, Grafana, Splunk

Responsibilities

Design scalable infrastructure solutions, build load/chaos/synthetic testing software, maintain automation tools for repetitive tasks, manage CPU/storage/GPU/network lifecycle platforms, implement fault-tolerant design patterns, develop and maintain SLOs/SLIs, participate in on-call rotation for critical incidents

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗