Distributed AI Support Engineer
Core
Provide first-line support, debugging, and operations for AI/LLM workloads on high-performance computing (HPC) infrastructure, enabling researchers and industry teams to train and run large-scale models.
Role type
Distributed AI Support Engineer (IC)
Builds
AI/LLM workflows, containerized environments, and support documentation for HPC and cloud stacks.
Domain
High-Performance Computing (HPC) and Artificial Intelligence
Deliverable
production ML models | infrastructure
Required skills
Python programming, PyTorch, distributed training frameworks (DDP, FSDP, DeepSpeed), GPU debugging, container management (Apptainer/Singularity), HPC job scheduling (Slurm), LLM inference frameworks (vLLM, Ray)
Preferred skills
LLM fine-tuning and quantization (QLoRA, PEFT), profiling tools (NVIDIA Nsight, TensorBoard), data I/O optimization, user training and documentation
Technologies
PyTorch, TensorFlow, Hugging Face Transformers, DeepSpeed, vLLM, Ray, Slurm, Apptainer, NVIDIA CUDA, NCCL, PyTorch Profiler, MLflow, Weights & Biases
Responsibilities
Triage and diagnose AI/HPC job failures, support users in writing and debugging multi-GPU job scripts, maintain and test shared AI/LLM software stacks, profile and tune distributed training workloads, advise on data storage and I/O bottlenecks, monitor usage metrics and prepare technical reports, develop documentation and training materials
Seniority
Mid-level, hands-on IC