CareerPlanSign in

Staff Site Reliability Engineer, Core AI Infrastructure

US - Remote Zone 1 (Job Requisitions Only)🌐 Remote💼 Full-time💰 $218,025–$218,025🗓 2026-06-08 → 2026-09-26

Core

Own the reliability, monitoring, and incident response lifecycle for critical AI infrastructure services, ensuring systems are resilient, observable, and secure at scale.

Role type

Staff Site Reliability Engineer (AI Infrastructure)

Builds

AI infrastructure services, CI/CD frameworks, and internal AI product applications

Domain

Cloud Infrastructure / AI Operations

Deliverable

production ML models | infrastructure

Required skills

Cloud infrastructure automation, Container orchestration, Incident response, Infrastructure as Code, Scripting/Programming, CI/CD pipelines

Preferred skills

Linux administration, Log aggregation, Network security, Regulated environment experience

Technologies

AWS, Terraform, Ansible, Chef, Puppet, Salt, Docker, Kubernetes, Go, Python, Bash, Ruby, Git

Responsibilities

Own reliability, monitoring, and incident response lifecycle for AI infrastructure services; Build automation and tooling to streamline operational IT workflows; Partner with Infrastructure and Security teams to extend CI/CD frameworks; Strengthen observability and documentation standards; Develop full-stack applications for internal AI products and infrastructure

Seniority

Staff, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.