CareerPlanSign in

Site Reliability Engineering (SRE ) lead

Atlanta, GA, US💼 Full-time💰 $111,605–$111,605🗓 2026-09-08 → 2026-09-27

Core

Lead troubleshooting, root cause analysis, and automation for complex production incidents and cloud platform reliability.

Role type

Senior SRE Lead (IC with team leadership)

Builds

Production-ready cloud infrastructure, automated operational runbooks, and self-healing capabilities

Domain

Financial Services / Cloud Infrastructure & Distributed Systems

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Incident management, Root cause analysis (RCA), Infrastructure as Code (IaC), CI/CD pipeline design, Cloud platform operations, Team leadership, Operational metrics analysis

Preferred skills

Distributed systems architecture, Automation framework development, Stakeholder management, Change management methodologies

Technologies

AWS, Azure, Kubernetes, Docker, Python, PowerShell, Shell, GitHub Actions, Azure DevOps, Jenkins, GitLab, Datadog, Splunk, Dynatrace, Grafana, Prometheus, CloudWatch, Azure Monitor, OpenTelemetry, ServiceNow, Jira, Terraform, Ansible, SQL

Responsibilities

Lead troubleshooting and resolution of complex production incidents; Conduct comprehensive root cause analysis and implement corrective actions; Design and enhance monitoring, observability, and alerting systems; Drive automation initiatives using scripting and IaC; Partner with engineering teams to remediate recurring reliability issues; Serve as Incident Commander during major outages; Provide leadership, coaching, and mentoring for SRE and production support engineers

Seniority

Senior, hands-on IC with team leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.