Senior Performance Engineer
Core
Build and evolve systems for performance analysis, telemetry, and optimization of large-scale GPU and CPU clusters for AI and HPC.
Role type
Senior IC performance engineer (AI/HPC infrastructure)
Builds
Performance analysis, benchmarking, and diagnostic tools for GPU/CPU clusters
Domain
High-performance computing (HPC) and Artificial Intelligence (AI)
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Performance analysis, systems engineering, high-performance networking (RDMA, MPI, NCCL), system performance metrics, hardware/firmware telemetry, Linux environments
Preferred skills
CUDA internals, congestion control algorithms, CPU/GPU architecture, deep learning frameworks (PyTorch, TensorFlow), cloud platforms, Python, Bash, C/C++
Technologies
NCCL, RDMA, MPI, RoCE, CUDA, PyTorch, TensorFlow, Linux
Responsibilities
Profile and benchmark AI/HPC workloads on GPU/CPU clusters, explore performance characteristics of high-performance networking, identify bottlenecks across networking/compute/memory/architecture, develop performance analysis tools, define performance test plans, collaborate with hardware/firmware/networking teams, support telemetry collection and data refinement
Seniority
Senior, hands-on IC
