TIONE-AI Infra 后台研发工程师
Core
Design and build a cloud-native large model inference platform on Kubernetes, focusing on scheduling, resource management, and core plugins to support efficient deployment of HunYuan and third-party models.
Role type
Senior IC Kubernetes infrastructure engineer (LLM inference)
Builds
High-elasticity, multi-tenant, heterogeneous compute unified scheduling platform for large model inference
Domain
Cloud-native infrastructure / AI inference / Distributed systems
Deliverable
production ML models
Required skills
Go, Kubernetes core mechanisms (Scheduler, CRD, Operator, Device Plugin, CSI, CNI), GPU/NPU resource management, high-performance networking (RDMA/RoCE), distributed system design
Preferred skills
vLLM, TensorRT-LLM, SGLang integration, open-source community contributions (Kubernetes, Volcano, KubeFlow), domestic AI chip adaptation (Ascend, Hygon, Tianxuan)
Technologies
Kubernetes, Go, C++, Python, RDMA, RoCE, CSI, CNI, Operator, controller-runtime
Responsibilities
Design control and data plane architecture for the inference platform; customize K8s scheduler for topology-aware GPU/NPU scheduling and resource pooling; develop core CRDs and Controllers for lifecycle management; implement device plugins for heterogeneous compute discovery; integrate and customize CNI for high-performance pod networking.
Seniority
Senior, hands-on IC