Senior Software Engineer, AI Resiliency
Core
Lead the development of AI software resiliency features for the world's most powerful AI supercomputers, ensuring robustness and reliability at a scale of 100,000+ GPUs.
Role type
Senior IC software engineer (AI infrastructure)
Builds
Production-grade AI software resiliency features for large-scale GPU clusters
Domain
AI infrastructure / High-Performance Computing (HPC)
Deliverable
production ML models
Required skills
C++, Python, distributed systems, parallel programming, fault tolerance, debugging, profiling, AI frameworks (PyTorch, JAX/XLA, TensorFlow), CUDA, NCCL, MPI
Preferred skills
Model training experience, checkpointing strategies, error mitigation, low-level performance tuning
Technologies
C++, Python, PyTorch, JAX, XLA, TensorFlow, CUDA, NCCL, MPI, gdb, perf, valgrind, NVIDIA Nsight
Responsibilities
Implement and optimize software features for AI system reliability (checkpoint-recovery, error detection, isolation, straggler/hang detection); Contribute to large-scale distributed systems with high-performance code; Work on AI system error handling and silent data corruption detection; Develop and implement tests for robustness and scalability; Assist in debugging and performance tuning large-scale AI workloads in cloud and HPC environments
Seniority
Senior, hands-on IC