CareerPlanGet AI match score →

Principal Site Reliability Engineer

2 Locations💼 Full-time🗓 2026-07-09 → 2026-07-31

Core

Designing, building, and optimizing cloud infrastructure and deployment systems to ensure scalability, security, and operational efficiency.

Role type

Principal Site Reliability Engineer (IC)

Builds

Cloud infrastructure, deployment systems, internal tools, and CI/CD pipelines

Domain

Cloud infrastructure, DevOps, Site Reliability Engineering

Deliverable

production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work

Required skills

Linux systems administration, Infrastructure as Code (Terraform, Packer, Ansible), Python programming, Golang programming, containerization (Docker), orchestration (AWS EKS, GCP GKE), GitOps workflows, CI/CD pipeline management, security program implementation, monitoring and logging (Prometheus, Grafana, ELK), distributed systems management (Apache Kafka, Cassandra)

Preferred skills

Open-source project contributions, security engineering background

Technologies

AWS, GCP, Terraform, Packer, Ansible, Python, Golang, Docker, Kubernetes, FluxCD, Jenkins, Prometheus, Grafana, ELK, Apache Kafka, Cassandra

Responsibilities

Enhance Infrastructure as Code and enforce best practices; Optimize cloud infrastructure for scalability, security, and cost-effectiveness; Develop internal tools to support cloud platform operations; Improve CI/CD pipelines and deployment workflows; Address container image vulnerabilities and standardize remediation processes; Build Amazon Machine Images aligned with security benchmarks; Strengthen monitoring, alerting, and observability; Troubleshoot complex production issues; Fine-tune distributed systems; Collaborate with development, security, and operations teams

Seniority

Principal, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗