High Performance Computing Engineer
Core
Design, operate, and maintain large-scale HPC environments, including deployment, configuration, and day-to-day operation of HPC schedulers and core HPC domains like GPU compute, high-performance storage, and networking.
Role type
Senior IC High Performance Computing Engineer
Builds
Large-scale HPC environments, high-scale training clusters, and scalable services on public cloud infrastructure
Domain
High Performance Computing, Cloud Infrastructure, AI/ML Training
Deliverable
infrastructure
Required skills
HPC scheduler management (SLURM, Kubernetes), GPU compute systems, high-performance storage, networking, Bash scripting, Python automation, public cloud infrastructure (Azure, AWS, GCP), LLM training clusters, NVIDIA InfiniBand, ML framework deployment and scaling
Preferred skills
Experience with NVIDIA H100/GB200, AI platforms and APIs, Machine Learning frameworks for language learning models
Technologies
SLURM, Kubernetes, Bash, Python, Azure, AWS, GCP, NVIDIA InfiniBand, Ray, NVIDIA H100, NVIDIA GB200
Responsibilities
Deploy, configure, and operate HPC schedulers; maintain and tune core HPC domains (GPU, storage, networking); develop automation and tooling for cluster reliability and observability; troubleshoot cluster usage issues and resolve failed jobs; support researchers and engineers with workload optimization
Seniority
Senior, hands-on IC