Lead Site Reliability Engineer
Core
Lead Site Reliability Engineer responsible for improving reliability and stability of web hosting platforms using data-driven analysis and SRE practices.
Role type
Senior IC Site Reliability Engineer
Builds
Reliable and stable web hosting platforms
Domain
Financial services infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site reliability engineering (SRE) practices, observability, monitoring, alerting, telemetry, service level objectives (SLOs), error budgets, incident triage, CI/CD quality checks, test automation, Python or Ansible, enterprise architecture, security, scalability, performance optimization
Preferred skills
AI-assisted operational workflows, SDLC knowledge, mentoring, knowledge sharing
Technologies
Grafana, Dynatrace, Prometheus, CloudWatch, Splunk, AWS, Azure, CI/CD tools, Ansible, Python
Responsibilities
Champion site reliability culture and practices; lead efforts to improve reliability and stability of web hosting platforms; define service level indicators and objectives aligned with customer needs; identify and remove technology bottlenecks; use enterprise-authorized AI capabilities to speed up incident triage and post-incident analysis; partner with technical specialists to solve complex issues; drive reuse-first adoption of AI-assisted reliability workflows across the SDLC; document and share knowledge across the organization
Seniority
Senior, hands-on IC