Lead Staff Systems Reliability Engineer (Linux & Distributed Systems)
Core
Lead a team to build, maintain, and optimize Linux-based distributed systems (Aerospike, Kafka, MongoDB) for real-time digital advertising workloads at internet scale.
Role type
Lead Staff Systems Reliability Engineer (Linux & Distributed Systems)
Builds
Global data-driven advertising platform handling massive datasets with sub-millisecond latency
Domain
Digital Advertising / Distributed Systems / High-Performance Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux OS, Leadership & Mentoring, Troubleshooting, Performance Tuning, Distributed Systems Architecture, Hardware Benchmarking, NoSQL Databases, Infrastructure Automation
Preferred skills
Physical Hardware Internals, Testing & Tuning, Kubernetes, Python/Ruby/Rust/Bash/Golang/C#, Prometheus
Technologies
Aerospike, MongoDB, Kafka, NVMe, Linux, Kubernetes, Prometheus, Ansible/PyInfra/Chef
Responsibilities
Lead team planning and work streams for global infrastructure; Build and improve infrastructure automation for stateful systems; Own operations for Linux-based systems; Benchmark and analyze next-gen hardware offerings; Participate in on-call rotation
Seniority
Lead Staff, hands-on IC with team leadership
