Staff+ Software Engineer, Infrastructure (Distributed Systems)
Core
Design, build, and operate large-scale distributed systems and infrastructure that train, serve, and secure AI models, including data pipelines, Kubernetes clusters, databases, and developer tooling.
Role type
Staff+ Software Engineer (Infrastructure)
Builds
Production-grade distributed systems, data pipelines, Kubernetes clusters, databases, observability stacks, and developer tooling for AI model training and serving.
Domain
AI Infrastructure / Distributed Systems / Cloud Computing
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed systems design, architectural decision-making, cloud infrastructure (AWS/GCP), Kubernetes, infrastructure-as-code, Python/Rust/Go/Java, incident response, technical mentorship.
Preferred skills
Machine learning infrastructure (GPUs/TPUs/Trainium), low-level systems (Linux kernel/eBPF), security/privacy engineering, technical leadership.
Technologies
Kubernetes, AWS, GCP, Python, Rust, Go, Java, NCCL, Linux kernel, eBPF.
Responsibilities
Independently scope and lead complex, multi-month infrastructure projects; make architectural decisions shaping the foundation for other teams; drive alignment on technical direction across multiple teams; partner with research and product teams to translate compute needs into technical designs; take ownership of reliability, scalability, and security; set technical strategy and standards; build operational processes (incident response, postmortems); mentor engineers.
Seniority
Staff+, hands-on IC with strategic scope and mentorship.
