Principal Software Engineer, AI Infra Management and Ops
Core
Design, develop, and evolve cloud-native services, distributed systems, and platform capabilities supporting large-scale production environments, including AI infrastructure and HPC platforms.
Role type
Principal Software Engineer (AI Infra Management and Ops)
Builds
Cloud-native services, distributed systems, platform capabilities, AI infrastructure, HPC platforms, bare-metal infrastructure, GPU environments, and large-scale distributed computing systems.
Domain
Cloud infrastructure, High Performance Computing (HPC), AI infrastructure, Datacenter architecture, Networking
Deliverable
production ML models | infrastructure
Required skills
Cloud-native architecture, distributed systems design, platform engineering, CI/CD, Infrastructure as Code (IaC), observability, reliability engineering, scalability design, technical mentorship, cross-organizational collaboration, large-scale system operations, hardware lifecycle management, fleet operations, network topology design, accelerated computing platforms, bare-metal infrastructure management, programming in Go, Rust, C#, Java, or Python
Preferred skills
Experience with InfiniBand fabrics, NVLink, NVSwitch, Ethernet fabrics, RDMA, SmartNICs, DPUs, rack-scale systems, datacenter networking technologies
Technologies
Go, Rust, C#, Java, Python, InfiniBand, NVLink, NVSwitch, Ethernet, RDMA, SmartNICs, DPUs
Responsibilities
Lead design and development of cloud-native services and distributed systems; Collaborate on datacenter architecture and networking solutions; Guide development of platform engineering capabilities for deployment automation and lifecycle management; Contribute to design and operation of AI infrastructure and HPC platforms; Foster technical excellence through architecture collaboration and mentorship
Seniority
Principal, hands-on IC with strategic influence