Site Reliability Engineer (SRE)
Core
Lead incident response and reliability engineering for Tmap Mobility's enterprise services, ensuring infrastructure stability and operational maturity.
Role type
Senior Site Reliability Engineer (SRE)
Builds
AWS multi-account + EKS production infrastructure, observability pipelines, and standardized deployment/change management processes.
Domain
Cloud Infrastructure (AWS, EKS) and Observability
Deliverable
production ML models | infrastructure
Required skills
Incident management and RCA, AWS operations, Observability platform usage, Server/Network troubleshooting, System interdependency analysis, Chaos Engineering planning
Preferred skills
High-traffic service operations, SLI/SLO/Error Budget management, Java/Spring performance analysis, LLM/Agent development and operations, Post-mortem culture establishment, Deployment/Change management automation
Technologies
AWS, EKS, Datadog, Claude Code, Codex
Responsibilities
Lead first-line response and RCA for enterprise service incidents, Manage post-mortems and prevent recurrence action items, Establish and standardize on-call/incident response systems, Define SLI/SLOs and quantify reliability metrics, Manage service operational maturity and production checklists, Standardize deployment/change management processes including impact analysis and risk review, Plan chaos engineering training and risk diagnostics
Seniority
Senior, hands-on IC