CareerPlanSign in

SRE AI高级工程师-基础架构

杭州💼 Full-time🗓 2026-09-28

Core

Manage and ensure consistency of massive high-performance GPU/XPU clusters for large model training, online inference, search, and recommendation training.

Role type

Senior SRE (AI Infrastructure)

Builds

Production GPU/XPU clusters for AI workloads

Domain

AI Infrastructure / High-Performance Computing

Deliverable

infrastructure

Required skills

GPU/XPU resource management, cluster scheduling, system architecture design, Go/Python/Java/C++, production troubleshooting, performance tuning, distributed training frameworks (TensorFlow/PyTorch)

Preferred skills

NVIDIA H100/A100 optimization, Ascend/XPU experience, global collaboration, project management

Technologies

NVIDIA H100, A100, Ascend, TensorFlow, PyTorch

Responsibilities

Deliver and guarantee consistency of massive GPU/XPU resources across training and inference scenarios; build automation to ensure stability and resource availability; isolate faulty GPU resources; design and manage the lifecycle of production clusters and services.

Sourced via bytedance · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.