Sr Site Reliability Engineer
Core
Strengthen the reliability and resilience of distributed payment systems by reducing manual effort and preventing operational incidents.
Role type
Senior Site Reliability Engineer (Payments Infrastructure)
Builds
T-Mobile payment platforms and distributed systems
Domain
Fintech / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Kubernetes, Python, Bash, CI/CD, Infrastructure as Code, Application Monitoring, Capacity Planning, Incident Management, Scripting, System Reliability
Preferred skills
NoSQL databases (Cassandra), SQL databases, Root cause analysis, Agile methodologies
Technologies
AWS, Kubernetes, Cassandra, SQL, Python, Bash
Responsibilities
Enhance system reliability by identifying issues and implementing preventive measures; Automate processes to accelerate software development and deployment; Conduct root cause analysis to prevent incident recurrence; Apply programming and scripting expertise to improve system robustness.
Seniority
Senior, hands-on IC