Production Engineer
Core
Drive an automation-first culture and ensure reliability for a global platform processing 200+ billion transactions daily across tens of millions of enterprise users.
Role type
Senior Production Engineer (Cloud Infrastructure & Operations)
Builds
Globally distributed, multi-cloud infrastructure (AWS, GCP, bare-metal) and self-healing systems
Domain
Cybersecurity / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, Go, C/C++, Linux/RHEL systems, networking protocols, distributed architecture, incident management, ITIL frameworks
Preferred skills
Infrastructure-as-Code (Ansible, Terraform, Helm, Temporal), chaos engineering, disaster recovery planning, global routing (BGP), traffic tunneling (GRE, IPSec), L7 proxy architectures (HAProxy), DNS at scale
Technologies
AWS, GCP, Prometheus, Grafana, OpenTelemetry, Python, Go, C/C++, Ansible, Terraform, Helm, Temporal, HAProxy
Responsibilities
Implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments; Write code to eliminate manual toil and build self-healing systems; Implement and maintain sophisticated observability, define SLIs/SLOs, and establish error budgets; Act as a lead Incident Commander, develop response playbooks, and conduct post-incident analyses; Partner with Engineering teams to conduct operability reviews
Seniority
Mid-level (1-3 years experience), hands-on IC