Lead Site Reliability Engineer
Core
Lead the production environment for distributed software applications, ensuring health, security, and availability through monitoring, tooling, and operational support.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production-grade monitoring tools, alerting systems, and automated infrastructure for large distributed applications.
Domain
Financial Services / Cloud Native / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, OpenShift, Azure, Java/SpringBoot, Python, SQL/No-SQL, Microservices, Container Orchestration, System Monitoring, Performance Tuning, Capacity Planning
Preferred skills
Mainframes & JCL, Multi-cloud/Hybrid environments, Docker, Splunk, Dynatrace, RUM, Grafana, OCP certification, GitHub
Technologies
Azure, OpenShift, Kubernetes, Docker, Splunk, Dynatrace, Grafana, RUM, Java, SpringBoot, Python, SQL, No-SQL
Responsibilities
Monitor system availability and holistic health; build tools for platform infrastructure management; debug production issues across the stack; drive tool creation for health monitoring and alerting; optimize system performance and reliability; analyze metrics for performance tuning and fault finding; participate in system design consulting and capacity planning; create sustainable systems through automation.
Seniority
Senior, hands-on IC with mentorship responsibilities