系统可用性与性能工程专家(GPU性能与稳定性工程师) - AI算力基础设施
Core
Design and develop fault diagnosis and risk governance tool platforms for million-scale data center servers and operating systems to ensure hardware and software stability.
Role type
Senior IC GPU performance and stability engineer
Builds
GPU performance benchmark suites, CPU/GPU observability tool platforms, and cluster-level GPU stability monitoring architecture
Domain
AI infrastructure, large-scale data centers, GPU clusters
Deliverable
production ML models
Required skills
GPU architecture and failure mode analysis, ECC/XID error code systems, NVLink/IB interconnect topology, CUDA programming, Python/Go system tool development, Prometheus/Grafana/ClickHouse observability stack
Preferred skills
AI business scenario operations, GPU cluster model training/inference support, PyTorch/SGLang/vLLM framework experience, multi-vendor GPU (NVIDIA/AMD/domestic) validation
Technologies
CUDA, Python, Go, Prometheus, Grafana, ClickHouse, NVLink, IB
Responsibilities
Collaborate with SRE teams to design stability tools; develop GPU performance benchmarks; build cluster-level stability monitoring; troubleshoot complex cross-layer availability and performance issues
Seniority
Senior, hands-on IC