Principal Software Engineer, AI Inference Cloud
Core
Define and build architecture for cloud-based AI inference services, developing highly available, scalable services for running AI inference workloads. (via careerplan.io/jobs/100399715616-principal-software-engineer-ai-inference-cloud-at-arm)
Role type
Principal Software Engineer (AI Inference Cloud)
Builds
Kubernetes controllers, platform capabilities for workload deployment/scheduling, and production inference services
Domain
Cloud Infrastructure / AI Inference
Deliverable
production ML models
Required skills
Kubernetes (controllers, operators, scheduling), Go/C++/Rust/Python, distributed systems, observability, incident response, architecture design, mentoring
Preferred skills
AI infrastructure, model serving, accelerator-backed workloads, PyTorch/Ray/vLLM/TensorRT-LLM, inference performance tuning
Technologies
Kubernetes, Go, C++, Rust, Python, PyTorch, Ray, vLLM, SGLang, TensorRT-LLM
Responsibilities
Define architecture for cloud-based AI inference services; Develop Kubernetes controllers and platform capabilities; Establish production practices for health validation and rollout; Lead production readiness reviews and resolve complex issues; Lead design reviews and mentor engineers
Seniority
Principal, hands-on IC & mentorship