百度百舸-机器学习平台研发工程师(GO方向)(J106189)
Core
Design and develop large-scale AI computing cluster infrastructure and products to support internal business and external customer needs.
Role type
Senior IC machine-learning platform engineer (Kubernetes/GPU)
Builds
Heterogeneous multi-core computing clusters, distributed training/inference systems, and end-to-end ML platform features for model development, training, deployment, and data engineering.
Domain
AI infrastructure, distributed systems, GPU computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes development, resource scheduling, container runtime, container networking, distributed training systems, distributed inference systems, GPU architecture
Preferred skills
Kubeflow, Volcano, PyTorch, Ray
Technologies
Kubernetes, GPU, PyTorch, Ray, Kubeflow, Volcano
Responsibilities
Design and develop large-scale AI computing cluster infrastructure; build heterogeneous multi-core computing clusters based on Kubernetes; optimize core capabilities like GPU cluster self-healing, resource scheduling, virtualization, and Serverless; develop large-scale distributed training and inference systems; build ML platforms supporting the full AI engineering lifecycle; improve service stability, performance, and scalability.