CareerPlanGet AI match score →

Senior AI Engineer

San Jose (CA)💼 Full-time💰 $209,000–$209,000🗓 2026-05-21 → 2026-07-31

Core

Design, build, and maintain the infrastructure and toolchains for large-scale distributed training and inference of Machine Learning models, specifically focusing on GPU clusters and LLMs.

Role type

Senior AI Infrastructure Engineer (LLM Training & Inference)

Builds

High-performance GPU training clusters, model serving architecture, and automated ML development pipelines.

Domain

Artificial Intelligence / Large Language Models / High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Distributed training systems (Horovod, DeepSpeed, PyTorch Distributed, Ray), Tensor/model parallelism, Cloud-native infrastructure (Kubernetes, Docker, Slurm), Systems programming (Python, Bash, C++, Go), Performance profiling and optimization, Automation (Prometheus, Grafana, Weights & Biases)

Preferred skills

Edge device deployment, Serverless architectures, Auto-scaling strategies, Model caching mechanisms

Technologies

Horovod, DeepSpeed, PyTorch Distributed, Ray, Kubernetes, Docker, Slurm, Prometheus, Grafana, Weights & Biases, CUDA, Triton, NCCL, gRPC, DALI, tf.data

Responsibilities

Develop and maintain high-performance LLM training GPU infrastructure and clusters; Optimize GPU utilization and implement fault-tolerant distributed training strategies; Build automated pipelines for data preprocessing, feature engineering, and model deployment; Design systems for dynamic resource allocation and model loading/unloading for inference; Develop dashboards and alerting mechanisms for real-time model performance monitoring; Provide technical support and troubleshooting for training and inference workloads.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗