Site Reliability Engineer, Associate
Core
Design and implement end-to-end monitoring solutions, drive capacity management and demand forecasting, and lead root cause investigations for production incidents to ensure reliability of full stack applications.
Role type
Associate Site Reliability Engineer (SRE)
Builds
Production ML models | product features | dashboards & analysis
Domain
Financial services / Infrastructure reliability
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Java (object-oriented programming), CI/CD practices, Telemetry solutions (Log/Performance monitoring, Grafana), Root cause analysis, Capacity management, Demand forecasting, Agile/Scrum methodologies, Troubleshooting performance issues
Preferred skills
Automated configuration management, AI/ML for problem solving, Scripting languages (Perl, Python), GIT
Technologies
Java, Grafana, GIT, Perl, Python
Responsibilities
Design and implement end-to-end monitoring solutions for Application and Infrastructure components; Drive the engineering of capacity management and demand forecasting solutions; Act as a culture carrier and leader, passing on SRE knowledge and best practices; Drive detailed root cause investigations for production incidents; Create/coordinate retros for significant incidents; Add custom Telemetry metrics to the code base of in-scope Applications
Seniority
Associate, hands-on IC