HPC System Engineer
Core
Benchmarking GPU platforms for machine learning and AI workloads to enable data-driven decisions for platform optimization and hardware development.
Role type
Systems Engineer (Cloudmeter)
Builds
GPU-based hardware evaluation and performance analysis for AI/ML frameworks
Domain
Cloud infrastructure, AI/ML, GPU hardware
Deliverable
production ML models
Required skills
Unix/Linux, Python, Bash, GPU stack (CUDA, NCCL, drivers), system troubleshooting, containerized environments (Docker, Kubernetes)
Preferred skills
Deep learning frameworks (PyTorch, JAX, vLLM, Tensort-LLM), job schedulers (Slurm, Volcano)
Technologies
CUDA, NCCL, Docker, Kubernetes, Slurm, Volcano, PyTorch, JAX, vLLM, Tensort-LLM
Responsibilities
Profile and analyze GPU performance at the system and kernel level; Evaluate and compare GPU performance across different platforms, architectures, and software stacks; Perform acceptance testing for new GPU clusters; Perform experiments across diverse GPU system configurations to assess interconnect strategies and system-level optimizations
Seniority
Individual Contributor