HPC Systems Engineer
Core
Design, integrate, and deliver cohesive end-to-end HPC and cloud infrastructure platforms supporting AI and scientific research workloads.
Role type
Senior IC HPC Systems Engineer
Builds
Large-scale AI and HPC deployment platforms integrating compute, storage, networking, and data center infrastructure
Domain
High-Performance Computing (HPC) and Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
HPC infrastructure design, distributed systems architecture, Kubernetes, data center infrastructure, infrastructure automation, system-level dependency analysis, cross-functional technical leadership
Preferred skills
GPU compute expertise (NVIDIA H100/H200), high-speed fabrics (InfiniBand, RoCEv2), scale-out storage systems (VAST Data, WekaFS, Lustre, GPFS), CI/CD pipeline design
Technologies
Kubernetes, Terraform, Ansible, InfiniBand, NVIDIA GPUs, Slurm, VAST Data, WekaFS, Lustre, GPFS
Responsibilities
Drive cross-domain technical alignment across compute, storage, networking, and data center teams; Lead integrated design reviews for new HPC deployments; Identify technical dependencies and integration risks; Define end-to-end architecture and engineering standards; Ensure designs meet performance and resiliency requirements; Develop system-level engineering specifications; Participate in failure analysis and operational readiness assessments; Collaborate with automation teams to improve deployment consistency; Serve as technical escalation point for cross-team integration issues
Seniority
Senior, hands-on IC