Principal Software Engineer, AI Compute Infrastructure
Core
Design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference.
Role type
Principal Software Engineer, AI Compute Infrastructure
Builds
Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management systems. (via careerplan.io/jobs/100399715872-principal-software-engineer-ai-compute-infrastructure-at-arm)
Domain
Cloud infrastructure, HPC, distributed systems, AI/ML hardware
Deliverable
infrastructure
Required skills
Kubernetes, Go/Python, Linux, networking, storage, GPU/accelerator support, distributed ML workloads, complex system troubleshooting
Preferred skills
Kubernetes scheduling/operators, NVIDIA technologies (CUDA, NVLink, NCCL), Terraform, Argo CD, Helm, Prometheus, PyTorch, Ray, vLLM, TensorRT-LLM
Responsibilities
Build and operate Kubernetes clusters with improved workload scheduling and capacity management; Enable new CPU and GPU systems by integrating drivers and monitoring; Investigate performance and reliability issues across infrastructure and hardware; Partner with AI teams to automate cluster provisioning and maintenance.
Seniority
Principal, hands-on IC with strategic guidance