SRE Engineer - San Francisco
Core
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating highly available, scalable cloud-native systems.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud-native systems for AI workloads
Domain
AI/ML infrastructure, Cloud platforms
Deliverable
production ML models | infrastructure
Required skills
Cloud platforms (AWS/GCP/Azure/OCI), Kubernetes, distributed systems, incident management, programming (Go/Python/Java), SLO/SLI definition
Preferred skills
AI/ML or data-intensive platform support, GPU cluster management, multi-region systems
Technologies
Kubernetes, AWS, GCP, Azure, OCI, Go, Python, Java
Responsibilities
Design and operate highly available, scalable cloud-native systems for AI workloads; Own production reliability, incident response, and on-call practices; Build observability (metrics, logs, tracing) and reliability automation; Define and manage SLOs, SLIs, and error budgets with engineering teams; Drive postmortems and reliability improvements across the platform; Partner with product and engineering teams on reliability design
Seniority
Senior, hands-on IC
