CareerPlanSign in

SRE AI高级工程师-基础架构

北京💼 Full-time🗓 2026-09-28

Core

Deliver and ensure consistency of massive high-performance GPU/XPU resources for large-scale model training, online inference, search, and recommendation training clusters.

Role type

Senior Site Reliability Engineer (AI Infrastructure)

Builds

Production-scale GPU/XPU clusters for AI workloads

Domain

AI Infrastructure / High-Performance Computing

Deliverable

infrastructure

Required skills

GPU/XPU resource management and scheduling, large-scale HPC cluster operations, system architecture design, Go/Python/Java/C++ development, production troubleshooting, performance tuning, distributed training frameworks (TensorFlow/PyTorch)

Preferred skills

NVIDIA H100/A100 optimization, systematic engineering thinking, product and engineering mindset, data structure and system design

Technologies

NVIDIA H100, NVIDIA A100, Ascend, TensorFlow, PyTorch

Responsibilities

Manage clusters for multi-scenario AI workloads, build automation for stability and resource efficiency, design and deploy production cluster services, isolate and resolve GPU faults

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.