Senior Site Reliability Engineer
Core
Senior SRE leading production stability, AI-augmented observability, and automation for Yahoo's US Commerce cloud-native infrastructure.
Role type
Senior Site Reliability Engineer (Infrastructure & Reliability)
Builds
Highly performant, reliable, and secure cloud-native infrastructure for Yahoo Commerce
Domain
Internet / Cloud Infrastructure / AI-augmented Observability
Deliverable
infrastructure
Required skills
AWS or GCP, Kubernetes, Docker, Python, Node.js, Go, Terraform, Ansible, Git, CI/CD pipelines, TCP/IP, Networking, Security best practices
Preferred skills
UNIX/Linux kernel-level troubleshooting, GitHub Actions, OpenTelemetry, Prometheus, Splunk, Grafana, Emerging AI tools
Technologies
AWS, GCP, Kubernetes, Docker, Terraform, Ansible, GitHub Actions, OpenTelemetry, Prometheus, Splunk, Grafana
Responsibilities
Lead management and optimization of production environments at scale; Design and implement AI-augmented observability frameworks; Drive operability roadmap via automation and CI/CD enhancements; Resolve complex system and network problems; Participate in senior-level on-call rotation and incident resolution
Seniority
Senior, hands-on IC with leadership scope