Site Reliability Engineering Lead
Core
Leading a team of Site Reliability Engineers to deliver reliable, scalable, and high-performing systems for the Brightmine Data Platform.
Role type
Senior IC Site Reliability Engineering Lead
Builds
The Brightmine Data Platform
Domain
Risk analytics and decisioning tools
Deliverable
production ML models | infrastructure
Required skills
Cloud platform management (AWS), Infrastructure as a Service (IaaS), DevOps practices, Site Reliability Engineering principles, Continuous Integration/Continuous Delivery (CI/CD), Distributed systems troubleshooting, Linux fundamentals, Monitoring and logging
Preferred skills
Disaster recovery planning, Infrastructure cost optimization, Incident management and Root Cause Analysis (RCA), Team coaching and performance management
Technologies
Amazon Web Services, Jenkins, GitLab, Terraform, Grafana, Prometheus
Responsibilities
Lead and manage SRE team performance, hiring, and development; Ensure alignment to SRE frameworks and best practices; Own prioritization of reliability engineering tasks; Lead incident management processes and post-mortems; Drive disaster recovery planning and production resilience testing; Support infrastructure cost analysis and optimization; Ensure team readiness for on-call support and incident response; Contribute to SRE capability building through training and coaching
Seniority
Senior, hands-on IC with people leadership