CareerPlanSign in

Software Engineer, Site Reliability (SRE)

San Francisco, CA💼 Full-time🗓 2025-10-21 → 2026-09-26

Core

Define and build the foundation of reliability, observability, and scalability across Sierra's AI-driven infrastructure, partnering with core engineering and product teams to ensure systems are highly available and built for growth.

Role type

Senior Site Reliability Engineer (SRE)

Builds

AI-driven infrastructure, observability stacks, and scalable cloud systems

Domain

Cloud Infrastructure & AI/ML Operations

Deliverable

production ML models | infrastructure

Required skills

Site Reliability Engineering, Infrastructure Engineering, Terraform, AWS services, Container Orchestration, Cloud Networking, Observability Systems, LLM Infrastructure, Incident Management, CI/CD Tooling

Preferred skills

LLM inference optimization, Fine-tuned model management, Large-scale model deployment, Startup environment experience, Self-healing infrastructure patterns

Technologies

Terraform, AWS, Prometheus, Grafana, Datadog

Responsibilities

Own the observability stack (monitoring, alerting, logging, tracing); Design reliable and scalable systems from day one; Implement secure cloud infrastructure; Improve reliability and scalability of LLM deployments; Lead improvements to deployment pipelines and incident management; Define SRE practices and influence engineering culture

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.