Senior Site Reliability Engineer (SRE) (all genders)
Core
Ensuring stability, high-performance, and scalability across the global application landscape by bridging Customer Operations, Software Development, and Infrastructure teams.
Role type
Senior Site Reliability Engineer (L3 Support)
Builds
Global e-commerce application landscape
Domain
E-commerce, Cloud-native (AWS)
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Monitoring and alerting implementation, Root Cause Analysis (RCA), Cloud-native environment management, Infrastructure as Code, CI/CD pipeline maintenance, Load and performance testing, Security vulnerability scanning, Compliance audit preparation
Preferred skills
Business Continuity planning, Disaster Recovery planning, Stakeholder collaboration
Technologies
Prometheus, Grafana, New Relic, AWS
Responsibilities
Implementing monitoring tools and dashboards for real-time system health tracking, Providing expert-level troubleshooting and performing root-cause analysis for complex incidents, Optimizing system performance and scalability under load, Identifying and mitigating security vulnerabilities, Ensuring operations meet legal and industry standards for audits
Seniority
Senior, hands-on IC