Software Engineer, Site Reliability (SRE)
Core
Define and build the foundation of reliability, observability, and scalability across Sierra's AI-driven infrastructure, partnering with core engineering and product teams to ensure systems are highly available and built for growth.
Role type
Senior Site Reliability Engineer (SRE)
Builds
AI-driven infrastructure, observability stacks, and scalable cloud systems
Domain
Cloud Infrastructure & AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Infrastructure Engineering, Terraform, AWS services, Container Orchestration, Cloud Networking, Observability Systems, LLM Infrastructure, Incident Management, CI/CD Tooling
Preferred skills
LLM inference optimization, Fine-tuned model management, Large-scale model deployment, Startup environment experience, Self-healing infrastructure patterns
Technologies
Terraform, AWS, Prometheus, Grafana, Datadog
Responsibilities
Own the observability stack (monitoring, alerting, logging, tracing); Design reliable and scalable systems from day one; Implement secure cloud infrastructure; Improve reliability and scalability of LLM deployments; Lead improvements to deployment pipelines and incident management; Define SRE practices and influence engineering culture
Seniority
Senior, hands-on IC