Software Engineer, Supercomputing
Core
Design, build, and operate GPU supercomputing environments for large-scale AI model training and inference.
Role type
Senior IC infrastructure engineer (supercomputing)
Builds
High-performance, reliable, and cost-efficient GPU clusters and orchestration software
Domain
AI infrastructure / High-performance computing
Deliverable
infrastructure
Required skills
Python or Rust, Kubernetes or Slurm, Linux systems administration, capacity planning, storage management, performance monitoring
Preferred skills
CUDA/NCCL, distributed training optimization, deep learning framework internals, infrastructure-as-code, networking
Technologies
Kubernetes, Slurm, Python, Rust, CUDA, NCCL, PyTorch, TensorFlow, JAX
Responsibilities
Operate and automate large GPU clusters, write software for cluster management, extend scheduling/orchestration systems, monitor operational metrics, build reliable storage paths, partner with researchers on scale runs
Seniority
Senior, hands-on IC