Site Reliability Engineer - Infrastructure Systems
Core
Manage and optimize Proton's global infrastructure of thousands of servers, ensuring high availability and resilience in adversarial network environments.
Role type
Senior Site Reliability Engineer (Infrastructure Systems)
Builds
Base platforms including Kubernetes, VM orchestration, and bare metal provisioning, plus critical services like DNS, DHCP, and monitoring.
Domain
Cloud Infrastructure & Network Security
Deliverable
production ML models | infrastructure
Required skills
Linux internals and kernel tuning, Kubernetes orchestration, Infrastructure as Code (Terraform, Ansible, Puppet), Python scripting, Distributed systems architecture, Observability (Prometheus, Grafana), Network security principles
Preferred skills
Bare metal provisioning, On-premise cloud solutions, Open-source contributions
Responsibilities
Oversee global server network expansion, Design automation for 99.95%+ uptime, Participate in on-call rotation for complex troubleshooting, Develop real-time monitoring and alerting systems, Engineer solutions for connectivity, stability, and scalability
Seniority
Senior, hands-on IC