Site Reliability Engineer Sr. Staff
Core
Design, build, and optimize cloud infrastructure and deployment systems to ensure scalability, security, and operational efficiency.
Role type
Sr. Staff Site Reliability Engineer (IC)
Builds
Cloud infrastructure, deployment systems, internal tools, and CI/CD pipelines
Domain
Cloud Infrastructure / DevOps / Site Reliability Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Linux systems administration, Cloud platforms (AWS, GCP), Infrastructure as Code (Terraform, Packer, Ansible), Python, Golang, Containerization (Docker), Orchestration (EKS, GKE), GitOps, CI/CD pipelines, Security compliance (CIS, STIG), Monitoring (Prometheus, Grafana, ELK), Distributed systems (Apache Kafka, Cassandra)
Preferred skills
Open-source contributions, Security engineering background
Technologies
AWS, GCP, Terraform, Packer, Ansible, Python, Golang, Docker, Kubernetes (EKS, GKE), FluxCD, Jenkins, Prometheus, Grafana, ELK, Apache Kafka, Cassandra
Responsibilities
Enhance Infrastructure as Code (IAC) and enforce best practices; Optimize cloud infrastructure for scalability, security, and cost-effectiveness; Develop internal tools to support cloud platform operations; Improve CI/CD pipelines and deployment workflows; Address container image vulnerabilities and standardize remediation processes; Build Amazon Machine Images (AMIs) aligned with CIS and STIG benchmarks; Strengthen monitoring, alerting, and observability; Troubleshoot complex production issues; Fine-tune distributed systems like Apache Kafka and Cassandra; Collaborate with development, security, and operations teams
Seniority
Sr. Staff, hands-on IC with significant independent judgment and team leadership