CareerPlanSign in

SRE Engineer - San Francisco

San Francisco, CA💼 Full-time🗓 2026-06-29 → 2026-09-26

Core

Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating highly available, scalable cloud-native systems.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Cloud-native systems for AI workloads

Domain

AI/ML infrastructure, Cloud platforms

Deliverable

production ML models | infrastructure

Required skills

Cloud platforms (AWS/GCP/Azure/OCI), Kubernetes, distributed systems, incident management, programming (Go/Python/Java), SLO/SLI definition

Preferred skills

AI/ML or data-intensive platform support, GPU cluster management, multi-region systems

Technologies

Kubernetes, AWS, GCP, Azure, OCI, Go, Python, Java

Responsibilities

Design and operate highly available, scalable cloud-native systems for AI workloads; Own production reliability, incident response, and on-call practices; Build observability (metrics, logs, tracing) and reliability automation; Define and manage SLOs, SLIs, and error budgets with engineering teams; Drive postmortems and reliability improvements across the platform; Partner with product and engineering teams on reliability design

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.