Software Engineer - Platform Infrastructure (Rust, C++)
Core
Design, build, and implement large-scale distributed systems powering one of the world's largest supercomputing clusters for AI training.
Role type
Senior IC systems engineer (distributed infrastructure)
Builds
Distributed systems, supercomputing clusters, and productivity tools for AI training
Domain
AI infrastructure, high-performance computing, systems programming
Deliverable
production ML models | infrastructure
Required skills
Systems programming (C/C++/Rust), Kubernetes cluster architecture, Linux kernel internals, performance profiling and optimization, container orchestration, distributed system networking
Preferred skills
Low-level debugging (kernel/OS), TCP/IP stack knowledge, observability in distributed systems, GitOps workflows, container runtime management
Technologies
Kubernetes, Docker, containerd, crio, Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, Helm, GitOps
Responsibilities
Design and implement large-scale distributed systems; profile, debug, and optimize performance across GPUs, Linux kernel, and networking; collaborate on hardware/software/algorithm co-design; maintain codebase for scalability; develop team productivity tools
Seniority
Senior, hands-on IC