百度百舸-机器学习平台研发(实习)(J105000)
Core
Design and develop large-scale AI computing cluster infrastructure and products to support internal business and external customer needs for training and inference tasks.
Role type
Machine Learning Platform Engineer (Intern)
Builds
Heterogeneous multi-core computing clusters, distributed training and inference systems, and end-to-end AI engineering workflows for large model development.
Domain
Cloud Computing / Artificial Intelligence / Distributed Systems
Deliverable
production ML models
Required skills
Python, Go, Java, C/C++, Kubernetes, Linux, MySQL, Network fundamentals, System fundamentals
Preferred skills
Large model inference experience, Serverless architecture experience
Responsibilities
Design and develop large-scale AI computing cluster infrastructure; Build heterogeneous multi-core computing clusters based on Kubernetes; Develop distributed training and inference systems; Support the full lifecycle of AI engineering including model development, training, deployment, and data engineering; Optimize service stability, performance, and scalability.