CareerPlanSign in

AI Infrastructure Engineer

Beijing🌐 Remote💼 Full-time🗓 2026-03-20 → 2026-09-25

Core

Build and scale AI inference infrastructure, managing the reliability, scalability, and operability of the model serving stack connecting research models to product systems.

Role type

Senior IC AI Infrastructure Engineer

Builds

Production-grade AI inference serving platform, GPU resource management systems, and automated O&M tools

Domain

Cloud-native AI infrastructure, GPU virtualization, distributed systems

Deliverable

production ML models

Required skills

Go or Python, Linux internals, networking, distributed systems, Kubernetes, Docker, microservices, inference systems, task scheduling, resource management

Preferred skills

GPU platforms, custom Kubernetes schedulers, Ray, model serving frameworks, distributed inference, GPU multiplexing (MIG, MPS, vGPU), SRE, observability, cloud cost optimization

Technologies

Kubernetes, Docker, Ray, MIG, MPS, vGPU

Responsibilities

Build and scale AI inference infrastructure from 0-1 to 1-N, design and optimize CPU/GPU resource management, drive production-level GPU scheduling and multiplexing, optimize inference pipeline performance, contribute to system stability and disaster recovery, explore AI-native infrastructure and automated O&M

Seniority

Mid-level, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.