CareerPlanSign in

Senior Principal AI Engineer

💼 Full-time🗓 2026-08-06 → 2026-09-26

Core

Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models) across GPU clusters to enable faster, larger, and more reliable model training.

Role type

Senior Principal AI Engineer (ML Infrastructure)

Builds

Scalable GPU cluster orchestration, distributed communication layers, and productionized large-model training pipelines.

Domain

Artificial Intelligence / Machine Learning Infrastructure / High-Performance Computing

Deliverable

production ML models

Required skills

Distributed systems engineering, GPU cluster management, PyTorch distributed training, parallelism strategies (data/tensor/pipeline), low-level GPU communication and networking, memory optimization (activation checkpointing, ZeRO), fault tolerance at scale

Preferred skills

Experience with large language models or foundation models, background in HPC

Technologies

Slurm, Kubernetes, Ray, RunAI, NCCL, RDMA, InfiniBand, NVLink, PyTorch Distributed, Megatron-LM, DeepSpeed

Responsibilities

Build and optimize GPU cluster orchestration; optimize and debug distributed communication; scale large-model training; apply advanced memory optimization techniques; partner with research and applied ML teams to productionize pipelines

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.