Staff Software Engineer, AI Inference Cloud
Core
Define and build architecture for cloud-based AI inference services, developing highly available, scalable Kubernetes orchestration and workload management systems.
Role type
Staff Software Engineer (AI Inference Cloud Platform)
Builds
Cloud-based AI inference services, Kubernetes controllers, and platform capabilities for workload deployment, scheduling, and lifecycle management.
Domain
Cloud Infrastructure / AI Systems (via careerplan.io/jobs/100399715408-staff-software-engineer-ai-inference-cloud-at-arm)
Deliverable
production ML models
Required skills
Kubernetes (controllers, operators, scheduling), Go/C++/Rust/Python, distributed systems, observability, incident response, architecture design, mentoring
Preferred skills
AI infrastructure, model serving, accelerator-backed workloads, PyTorch/Ray/vLLM/TensorRT-LLM, inference performance tuning
Technologies
Kubernetes, Go, C++, Rust, Python, PyTorch, Ray, vLLM, SGLang, TensorRT-LLM
Responsibilities
Define and build architecture for cloud-based AI inference services; Develop Kubernetes controllers and platform capabilities; Establish production practices for health validation and progressive rollout; Lead production readiness reviews and resolve complex infrastructure issues; Lead design reviews and mentor engineers
Seniority
Staff, hands-on IC with strategic influence