Lead Principal Site Reliability Engineer CSS
Core
Lead SRE responsible for maintaining high availability, reliability, and performance of mission-critical cloud infrastructure and enterprise applications.
Role type
Lead Principal Site Reliability Engineer
Builds
Automated cloud infrastructure, CI/CD pipelines, and resilient containerized workloads
Domain
Cloud Infrastructure / Financial Services
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Docker, Terraform, Python, Bash, Linux administration, CI/CD pipeline design, observability stack management, capacity planning, root cause analysis, disaster recovery planning
Preferred skills
Oracle Cloud Infrastructure (OCI), AWS, Azure, GCP, Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, New Relic, networking concepts (DRG, DNS, load balancing, routing, APIs)
Technologies
AWS, Azure, GCP, OCI, Kubernetes, Docker, Terraform, Python, Bash, Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Prometheus, Grafana, ELK, OpenSearch, Splunk, Datadog, New Relic, Linux, Git
Responsibilities
Maintain high availability and performance across enterprise applications and cloud infrastructure; Design and deliver automation to reduce manual effort; Build Infrastructure as Code and automation scripts; Develop and support CI/CD pipelines; Monitor applications and infrastructure through metrics, logs, and traces; Define and manage SLIs, SLOs, and error budgets; Participate in on-call rotations and respond to production incidents; Investigate incidents through root cause analysis; Carry out capacity planning and performance tuning; Support Kubernetes clusters and cloud-native applications; Partner with development teams to strengthen resiliency; Work with security teams to ensure compliance; Produce runbooks and technical documentation; Continuously enhance platform reliability through automation and best practices; Forecast infrastructure demand and respond to capacity needs; Identify resource gaps and support cost reduction efforts; Lead root cause analyses for incidents and maintenance activities; Deliver strategic health and performance reporting; Provide expert release notes and guidance on scale, capacity, security, and performance; Lead on-call shifts and resolve complex issues; Contribute to initiatives improving bottlenecks, deployments, and scalability; Share expertise on site reliability trends and shape best practices.
Seniority
Lead Principal, hands-on IC with strategic guidance