AI Infrastructure Engineer
Core
Build and scale AI inference infrastructure, managing the reliability, scalability, and operability of the model serving stack connecting research models to product systems.
Role type
Senior IC AI Infrastructure Engineer
Builds
Production-grade AI inference serving platform, GPU resource management systems, and automated O&M tools
Domain
Cloud-native AI infrastructure, GPU virtualization, distributed systems
Deliverable
production ML models
Required skills
Go or Python, Linux internals, networking, distributed systems, Kubernetes, Docker, microservices, inference systems, task scheduling, resource management
Preferred skills
GPU platforms, custom Kubernetes schedulers, Ray, model serving frameworks, distributed inference, GPU multiplexing (MIG, MPS, vGPU), SRE, observability, cloud cost optimization
Technologies
Kubernetes, Docker, Ray, MIG, MPS, vGPU
Responsibilities
Build and scale AI inference infrastructure from 0-1 to 1-N, design and optimize CPU/GPU resource management, drive production-level GPU scheduling and multiplexing, optimize inference pipeline performance, contribute to system stability and disaster recovery, explore AI-native infrastructure and automated O&M
Seniority
Mid-level, hands-on IC