CareerPlanSign in

AI Infra研发工程师 - 存储

北京💼 Full-time🗓 2026-09-28

Core

Develop and optimize distributed parallel strategies for deep learning frameworks to support efficient large-scale model training.

Role type

Senior IC distributed systems engineer (AI training infrastructure)

Builds

Distributed training frameworks and optimization layers for large language models

Domain

AI infrastructure / Distributed systems

Deliverable

production ML models

Required skills

Deep learning framework internals, distributed parallel strategies (data/tensor/pipeline parallelism), collective communication optimization (AllReduce, AllGather), C++, Python, Linux, algorithm and data structures

Preferred skills

Storage system development, CUDA programming, GPU parallel optimization

Responsibilities

Research and optimize distributed parallel strategies for deep learning frameworks; Optimize collective communication techniques to resolve bottlenecks in large-scale clusters; Collaborate on data scheduling and caching strategies for distributed training scenarios; Drive architecture iteration and troubleshoot issues in distributed training deployments

Sourced via bytedance · Listed on CareerPlan, which tracks 885,000+ jobs from 20+ sources.