Senior Infrastructure Engineer
Core
Design and develop core architectural improvements for GPU delivery, health monitoring, triage automation, and diagnostic services to support distributed AI, ML, and HPC workloads across thousands of GPUs.
Role type
Senior Infrastructure Engineer (GPU/AI/ML)
Builds
High-performance infrastructure for distributed AI, ML, and HPC workloads using RoCE and Infiniband
Domain
Cloud Infrastructure / AI & Machine Learning / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Java or Go, Python or Shell scripting, Git, Bitbucket, distributed systems, CI/CD pipelines, low-level systems troubleshooting
Preferred skills
Public cloud platforms (AWS, Azure, Oracle), multi-AD/AZ environments, large distributed systems, internal customer partnership
Technologies
AWS, Azure, Oracle, RoCE, Infiniband, Java, Go, Python, Shell, Git, Bitbucket, CI/CD
Responsibilities
Design and develop core architectural improvements for GPU delivery and diagnostic services; build systems supporting distributed AI, ML, and HPC workloads; debug and troubleshoot low-level stack issues; maintain simplicity and scalability in design and implementation; collaborate with Network and Data Center operations teams.
Seniority
Senior, hands-on IC
