ML Infrastructure Engineer
Core
Lead and support benchmarking of GPU platforms for machine learning and AI workloads, evaluating performance to enable data-driven decisions for platform optimization and hardware development.
Role type
ML Infrastructure Engineer (GPU Benchmarking & Optimization)
Builds
Full-stack AI cloud platform supporting developers and enterprises from data/model training to production deployment
Domain
Cloud Infrastructure / AI Hardware / GPU Systems
Deliverable
production ML models
Required skills
GPU performance profiling, deep learning frameworks (PyTorch, JAX, Megatron-LM, Tensort-LLM), CUDA/ROCm stack, neural network parallelism optimization, containerization (Docker, Kubernetes), performance visualization tooling
Preferred skills
LLM inference frameworks (vLLM, SGLang, TensorRT), Python performance profiling (Nsight, nvprof, perf), cloud ML platforms (AWS, GCP, Azure ML), open-source ML benchmarking contributions
Technologies
CUDA, ROCm, NCCL, Docker, Kubernetes, PyTorch, JAX, Megatron-LM, Tensort-LLM, vLLM, SGLang, TensorRT, Nsight, nvprof, perf, AWS, GCP, Azure ML
Responsibilities
Profile and analyze GPU performance at system and kernel level; Evaluate and compare GPU performance across platforms, architectures, and software stacks; Debug and optimize ML workloads to run efficiently on GPU hardware; Perform acceptance testing for new GPU clusters; Conduct experiments on GPU system configurations to assess interconnect strategies and system-level optimizations; Develop tools and dashboards to visualize performance metrics, bottlenecks, and trends
Seniority
Mid-Senior, hands-on IC