CareerPlanSign in

TIONE-AI Infra 后台研发工程师

Shanghai, China💼 Full-time🗓 2026-09-28

Core

Design and build a cloud-native large model inference platform on Kubernetes, focusing on scheduling, resource management, and core plugins to support efficient deployment of HunYuan and third-party models.

Role type

Senior IC Kubernetes infrastructure engineer (LLM inference)

Builds

High-elasticity, multi-tenant, heterogeneous compute unified scheduling platform for large model inference

Domain

Cloud-native infrastructure / AI inference / Distributed systems

Deliverable

production ML models

Required skills

Go, Kubernetes core mechanisms (Scheduler, CRD, Operator, Device Plugin, CSI, CNI), GPU/NPU resource management, high-performance networking (RDMA/RoCE), distributed system design

Preferred skills

vLLM, TensorRT-LLM, SGLang integration, open-source community contributions (Kubernetes, Volcano, KubeFlow), domestic AI chip adaptation (Ascend, Hygon, Tianxuan)

Technologies

Kubernetes, Go, C++, Python, RDMA, RoCE, CSI, CNI, Operator, controller-runtime

Responsibilities

Design control and data plane architecture for the inference platform; customize K8s scheduler for topology-aware GPU/NPU scheduling and resource pooling; develop core CRDs and Controllers for lifecycle management; implement device plugins for heterogeneous compute discovery; integrate and customize CNI for high-performance pod networking.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.