Staff, Site Reliability Engineer (Tech Infra)
Core
Build, operate, and scale large-scale e-commerce systems ensuring stability, performance, and high availability for customer-facing services.
Role type
Staff Site Reliability Engineer (Tech Infra)
Builds
Highly available, automated, and scalable e-commerce infrastructure and services
Domain
E-commerce / Distributed Systems / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Large-scale distributed systems architecture, UNIX/Linux system administration, Python/Java/Golang/Ruby programming, System/network troubleshooting, Cloud infrastructure management, CI/CD and Infrastructure as Code, Container orchestration, Observability tooling
Preferred skills
Large-scale web-based Java architecture and JVM tuning, Cloud monitoring certifications, Large-scale e-commerce platform experience
Technologies
AWS, Azure, Google Cloud Platform, Terraform, Docker, Kubernetes, Prometheus, Grafana, Elastic Stack, Datadog, New Relic
Responsibilities
Define and manage KPIs and SLOs for system availability and performance, Build and maintain incident management and disaster recovery automation, Establish best practices for monitoring and telemetry systems, Collaborate with product teams on scalable and operable design, Participate in 24x7 on-call rotation for rapid issue resolution
Seniority
Staff, hands-on IC with strategic influence