HPC Performance Engineer
Core
Design, develop, and optimize bare-metal HPC systems from POST through Kubernetes cluster integration, ensuring low-latency, high-throughput performance for AI workloads.
Role type
Senior HPC Performance Engineer (Systems)
Builds
Custom Linux kernels, OS images, virtualization stacks, and container runtimes for AI infrastructure
Domain
High-Performance Computing (HPC) / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux internals, MPI workloads, distributed system performance analysis, RoCE/InfiniBand/GPUDirect, HPC benchmarking (HPCC/HPL/MLPerf-HPC), Python automation, systems debugging
Preferred skills
Golang, QA/QE best practices, cloud environments, open-source development, machine learning
Technologies
Linux Kernel, Ubuntu, Prometheus, Victoria Metrics, Grafana, Docker, Kubernetes, KubeVirt, containerd, InfiniBand, RoCE, Nvidia GPUs
Responsibilities
Develop tools for systems performance baselines; maintain performance regression analysis testing automation; debug and tune fabric-level performance; develop telemetry for distributed clusters; triage and fix Linux performance issues; define OS requirements and system architecture for performance
Seniority
Senior, hands-on IC