Principal Site Reliability Engineer
Core
Design, deploy, and maintain scalable, secure applications and infrastructure in cloud or hybrid environments to ensure service robustness, scalability, and maintainability.
Role type
Principal Site Reliability Engineer
Builds
Highly available, reliable, deployable systems and operational processes
Domain
Transportation technology and defense capabilities
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud platform expertise (AWS, GCP, Azure), container orchestration (Docker, Kubernetes), scripting (PowerShell, Python, Go, Bash), Infrastructure-as-Code (Terraform), monitoring and observability, CI/CD pipelines, SLO/SLI management, incident response, capacity planning, disaster recovery design
Preferred skills
SRE-specific certifications, experience shaping and scaling SRE practices, team mentoring
Technologies
AWS, GCP, Azure, Docker, Kubernetes, Terraform, Jenkins, GitLab CI/CD, Argo CD, Python, Go, Bash, PowerShell
Responsibilities
Design and maintain scalable infrastructure; implement monitoring and alerting systems; automate operational tasks; lead incident response and RCA; conduct performance tuning and capacity planning; reduce manual toil in security compliance; support troubleshooting across environments; guide project rollouts; identify organization-wide reliability gaps
Seniority
Principal, hands-on IC with strategic influence