CareerPlanGet AI match score →

Software Engineer, Infrastructure Reliability

San Francisco💼 Full-time🗓 2026-03-19 → 2026-07-31

Core

Design, build, and operate reliable, performant, and secure distributed systems and cloud infrastructure to support global-scale AI research and products.

Role type

Senior IC Infrastructure Reliability Engineer

Builds

Core Distributed Systems, Databases, Observability, and Cloud Infrastructure platforms

Domain

AI/ML Infrastructure, Cloud Computing, Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

distributed systems design, Kubernetes orchestration, cloud infrastructure (AWS/GCP/Azure), Infrastructure as Code (Terraform), observability stack management, microservices architecture, Linux administration, performance optimization, incident response, automation scripting

Preferred skills

service mesh technologies, database technologies, networking expertise, security best practices in cloud environments

Technologies

Kubernetes, Terraform, Datadog, Prometheus, Grafana, Splunk, ELK stack, AWS, GCP, Azure

Responsibilities

Design and operate reliable systems used across engineering; identify and fix performance bottlenecks to enable scaling; resolve complex technical issues; improve automation and internal tooling; contribute to incident response and postmortems

Seniority

Senior, hands-on IC with leadership experience

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗