CareerPlanSign in

系统可用性与性能工程专家(GPU性能与稳定性工程师) - AI算力基础设施

北京💼 Full-time🗓 2026-09-28

Core

Design and develop fault diagnosis and risk governance tool platforms for million-scale data center servers and operating systems to ensure hardware and software stability.

Role type

Senior IC GPU performance and stability engineer

Builds

GPU performance benchmark suites, CPU/GPU observability tool platforms, and cluster-level GPU stability monitoring architecture

Domain

AI infrastructure, large-scale data centers, GPU clusters

Deliverable

production ML models

Required skills

GPU architecture and failure mode analysis, ECC/XID error code systems, NVLink/IB interconnect topology, CUDA programming, Python/Go system tool development, Prometheus/Grafana/ClickHouse observability stack

Preferred skills

AI business scenario operations, GPU cluster model training/inference support, PyTorch/SGLang/vLLM framework experience, multi-vendor GPU (NVIDIA/AMD/domestic) validation

Technologies

CUDA, Python, Go, Prometheus, Grafana, ClickHouse, NVLink, IB

Responsibilities

Collaborate with SRE teams to design stability tools; develop GPU performance benchmarks; build cluster-level stability monitoring; troubleshoot complex cross-layer availability and performance issues

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 833,000+ jobs from 20+ sources.