CareerPlanSign in

AI基础设施架构师/高级AI基础设施研究员-基础设施

北京💼 Full-time🗓 2026-09-28

Core

Design and evolve end-to-end architecture for ultra-scale, high-availability AI Factory clusters to support efficient pre-training and inference of billion/ trillion-parameter LLMs/VLMs.

Role type

Senior IC AI Infrastructure Architect / Researcher

Builds

Next-generation distributed scheduling systems, high-performance storage, and data pipelines for AI/ML workloads.

Domain

AI Infrastructure, Distributed Systems, High-Performance Computing

Deliverable

production ML models

Required skills

GPU cluster architecture, RDMA/RoCE networking, C++/Go/Rust system programming, distributed training frameworks (Megatron-LM, DeepSpeed), low-latency storage design, performance profiling

Preferred skills

First-principles problem solving, fault-tolerance mechanisms, topology-aware placement, compiler optimization (XLA/TVM)

Technologies

Kubernetes, NCCL, NVLink, InfiniBand, vLLM, Triton

Responsibilities

Design end-to-end AI Factory architecture for 100k+ GPU clusters; Implement distributed scheduling and resource management systems; Optimize full-stack AI/ML compute stack and eliminate system bottlenecks; Design high-throughput distributed storage and ETL pipelines.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 849,000+ jobs from 20+ sources.