MTS 1, Platform Reliability Engineer
Core
Own the reliability, operability, and evolution of internal engineering platforms, focusing on reducing toil and enabling safe AI-driven automation at scale.
Role type
Senior IC Platform Reliability Engineer (SRE)
Builds
Internal engineering platform, AI-powered operational workflows, and automated remediation systems
Domain
Ecommerce, Cloud Infrastructure, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
incident management, root cause analysis, CI/CD pipeline design, observability tooling, Linux administration, Java/Python/Go/Shell programming, security vulnerability remediation, SLO/SLI definition, automation framework development
Preferred skills
chaos engineering, self-healing system design, SRE frameworks, platform standardization
Technologies
Java, Python, Go, Shell, Linux, observability stacks, CI/CD tools
Responsibilities
Lead incident triage and post-mortems; build and operate AI-powered automation workflows; design guardrails for automated code/infrastructure changes; define and track reliability metrics; standardize deployment pipelines; remediate security vulnerabilities
Seniority
Senior, hands-on IC