CareerPlanSign in

Lead Principal Site Reliability Engineer CSS

USA💼 Full-time💰 $96,300–$96,300🗓 2026-09-04 → 2026-09-26

Core

Lead SRE responsible for maintaining high availability, reliability, and performance of mission-critical cloud infrastructure and enterprise applications.

Role type

Lead Principal Site Reliability Engineer

Builds

Automated cloud infrastructure, CI/CD pipelines, and resilient containerized workloads

Domain

Cloud Infrastructure / Financial Services

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Kubernetes, Docker, Terraform, Python, Bash, Linux administration, CI/CD pipeline design, observability stack management, capacity planning, root cause analysis, disaster recovery planning

Preferred skills

Oracle Cloud Infrastructure (OCI), AWS, Azure, GCP, Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, New Relic, networking concepts (DRG, DNS, load balancing, routing, APIs)

Technologies

AWS, Azure, GCP, OCI, Kubernetes, Docker, Terraform, Python, Bash, Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Prometheus, Grafana, ELK, OpenSearch, Splunk, Datadog, New Relic, Linux, Git

Responsibilities

Maintain high availability and performance across enterprise applications and cloud infrastructure; Design and deliver automation to reduce manual effort; Build Infrastructure as Code and automation scripts; Develop and support CI/CD pipelines; Monitor applications and infrastructure through metrics, logs, and traces; Define and manage SLIs, SLOs, and error budgets; Participate in on-call rotations and respond to production incidents; Investigate incidents through root cause analysis; Carry out capacity planning and performance tuning; Support Kubernetes clusters and cloud-native applications; Partner with development teams to strengthen resiliency; Work with security teams to ensure compliance; Produce runbooks and technical documentation; Continuously enhance platform reliability through automation and best practices; Forecast infrastructure demand and respond to capacity needs; Identify resource gaps and support cost reduction efforts; Lead root cause analyses for incidents and maintenance activities; Deliver strategic health and performance reporting; Provide expert release notes and guidance on scale, capacity, security, and performance; Lead on-call shifts and resolve complex issues; Contribute to initiatives improving bottlenecks, deployments, and scalability; Share expertise on site reliability trends and shape best practices.

Seniority

Lead Principal, hands-on IC with strategic guidance

Sourced via devitjobs · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.