Staff Production Engineer
Core
Drive an automation-first culture and ensure reliability for a global cloud platform processing 200+ billion daily transactions.
Role type
Staff Production Engineer (Cloud Infrastructure & Operations)
Builds
Globally distributed, multi-cloud infrastructure (AWS, GCP, bare-metal) and self-healing systems
Domain
Cybersecurity / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Python, Go, C/C++, Linux/RHEL systems, networking protocols, distributed architecture, incident management, ITIL frameworks
Preferred skills
Infrastructure-as-Code (Ansible, Terraform, Helm, Temporal), chaos engineering, disaster recovery planning, global routing (BGP), traffic tunneling (GRE, IPSec), L7 proxy architectures (HAProxy), DNS at scale
Technologies
AWS, GCP, Prometheus, Grafana, OpenTelemetry, Python, Go, C/C++, Ansible, Terraform, Helm, Temporal, HAProxy
Responsibilities
Implement highly available, scalable infrastructure; write code to eliminate manual toil and build self-healing systems; implement and maintain sophisticated observability and define SLIs/SLOs; act as lead Incident Commander and conduct post-incident analyses; partner with teams on operability reviews
Seniority
Staff, hands-on IC with strategic impact