Senior Performance Engineer- Pre-training(f/m/d)
Core
Engineer systems to train foundation models at scale by maximizing hardware utilization and training throughput on large-scale GPU clusters.
Role type
Senior IC performance engineer (pre-training)
Builds
High-throughput training infrastructure for large language models
Domain
AI/ML infrastructure, distributed systems, GPU computing
Deliverable
production ML models
Required skills
Python, PyTorch, CUDA programming, distributed systems, parallel computing, GPU microarchitecture, profiling tools (PyTorch Profiler, Nsight Systems, Nsight Compute), NCCL, MPI
Preferred skills
Distributed training frameworks (TorchTitan, Megatron-LM, DeepSpeed), low-precision training formats (MXFP4, MXFP8), NVIDIA Blackwell architecture, NVSHMEM, CUDA IPC
Technologies
PyTorch, CUDA, NCCL, MPI, Nsight Systems, Nsight Compute, NVIDIA Blackwell GPUs
Responsibilities
Profile training loops to identify system- and kernel-level bottlenecks; Configure and tune composite parallelism strategies (TP, DP, HSDP/FSDP, EP); Partner with AI Researchers to define model architectures for hardware efficiency
Seniority
Senior, hands-on IC