Senior Principal AI Engineer
Core
Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models) across GPU clusters to enable faster, larger, and more reliable model training.
Role type
Senior Principal AI Engineer (ML Infrastructure)
Builds
Scalable GPU cluster orchestration, distributed communication layers, and productionized large-model training pipelines.
Domain
Artificial Intelligence / Machine Learning Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed systems engineering, GPU cluster management, PyTorch distributed training, parallelism strategies (data/tensor/pipeline), low-level GPU communication and networking, memory optimization (activation checkpointing, ZeRO), fault tolerance at scale
Preferred skills
Experience with large language models or foundation models, background in HPC
Technologies
Slurm, Kubernetes, Ray, RunAI, NCCL, RDMA, InfiniBand, NVLink, PyTorch Distributed, Megatron-LM, DeepSpeed
Responsibilities
Build and optimize GPU cluster orchestration; optimize and debug distributed communication; scale large-model training; apply advanced memory optimization techniques; partner with research and applied ML teams to productionize pipelines
Seniority
Senior, hands-on IC
