CareerPlanSign in

Member of Technical Staff, AI Training Infrastructure

San Mateo💼 Full-time🗓 2026-07-31 → 2026-09-25

Core

Design, build, and optimize infrastructure for large-scale AI model training operations, including distributed pipelines and data storage.

Role type

Senior IC infrastructure engineer (AI training)

Builds

Scalable training pipelines, distributed training environments, and data storage solutions for LLMs and multimodal models

Domain

AI infrastructure, distributed systems, cloud computing

Deliverable

infrastructure

Required skills

distributed systems, ML infrastructure, PyTorch, cloud platforms (AWS/GCP/Azure), containerization, Kubernetes, Docker, distributed training techniques (data/model parallelism, FSDP)

Preferred skills

large language model training, multimodal AI systems, ML workflow orchestration, high-performance distributed computing, ML DevOps, open-source ML infrastructure contributions

Technologies

PyTorch, Kubernetes, Docker, AWS, GCP, Azure

Responsibilities

Design scalable infrastructure for large-scale model training; Develop and maintain distributed training pipelines for LLMs and multimodal models; Optimize training performance across GPUs, nodes, and data centers; Implement monitoring, logging, and debugging tools; Architect and maintain data storage solutions for large-scale training datasets; Automate infrastructure provisioning, scaling, and orchestration; Analyze and improve efficiency, scalability, and cost-effectiveness of training systems; Troubleshoot complex performance issues in distributed training environments

Seniority

Mid-level, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.