Principal Production Engineer
Core
Design and implement highly available, scalable infrastructure for a global platform processing 200+ billion transactions daily, driving an automation-first culture to reduce Mean Time to Mitigate.
Role type
Principal Production Engineer (Cloud Infrastructure & Operations)
Builds
Globally distributed, multi-cloud infrastructure (AWS, GCP, bare-metal) and self-healing systems
Domain
Cybersecurity / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Python, Go, C/C++, Linux/RHEL systems, networking protocols, distributed architecture, incident management, ITIL frameworks
Preferred skills
Infrastructure-as-Code (Ansible, Terraform, Helm, Temporal), chaos engineering, disaster recovery planning, global routing (BGP), traffic tunneling (GRE, IPSec), L7 proxy architectures (HAProxy), DNS at scale, OS networking stack internals
Responsibilities
Design and implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments; Drive an automation-first culture by writing code to eliminate manual toil and build self-healing systems; Implement and maintain sophisticated observability (Prometheus, Grafana, OpenTelemetry), define SLIs/SLOs, and establish error budgets; Act as a lead Incident Commander (TDO on-call), develop response playbooks, and conduct deep-dive post-incident analyses; Partner with Engineering and partner teams to conduct operability reviews
Seniority
Principal, hands-on IC with strategic vision