推理流量调度研发工程师-Data AML
Core
Design and optimize core architecture for trillion-level TPM traffic scheduling in AI model inference services.
Role type
Senior IC distributed systems engineer (inference traffic scheduling)
Builds
AI MaaS platform inference services on public cloud & Kubernetes
Domain
Cloud-native infrastructure, large-scale distributed systems, AI inference
Deliverable
production ML models
Required skills
Golang, C++, Python, Linux, data structures, algorithms, networking, operating systems
Preferred skills
Service Mesh, Kubernetes, Docker, public cloud storage/network/security, large model inference frameworks (vLLM, Triton)
Responsibilities
Develop and iterate core traffic scheduling architecture, design key inference subsystems, implement intelligent traffic routing and service governance, optimize resource efficiency and stability, tackle distributed system challenges
Seniority
Senior, hands-on IC