Member of Technical Staff, AI Training Infrastructure
Core
Design, build, and optimize infrastructure for large-scale AI model training operations, including distributed pipelines and data storage.
Role type
Senior IC infrastructure engineer (AI training)
Builds
Scalable training pipelines, distributed training environments, and data storage solutions for LLMs and multimodal models
Domain
AI infrastructure, distributed systems, cloud computing
Deliverable
infrastructure
Required skills
distributed systems, ML infrastructure, PyTorch, cloud platforms (AWS/GCP/Azure), containerization, Kubernetes, Docker, distributed training techniques (data/model parallelism, FSDP)
Preferred skills
large language model training, multimodal AI systems, ML workflow orchestration, high-performance distributed computing, ML DevOps, open-source ML infrastructure contributions
Technologies
PyTorch, Kubernetes, Docker, AWS, GCP, Azure
Responsibilities
Design scalable infrastructure for large-scale model training; Develop and maintain distributed training pipelines for LLMs and multimodal models; Optimize training performance across GPUs, nodes, and data centers; Implement monitoring, logging, and debugging tools; Architect and maintain data storage solutions for large-scale training datasets; Automate infrastructure provisioning, scaling, and orchestration; Analyze and improve efficiency, scalability, and cost-effectiveness of training systems; Troubleshoot complex performance issues in distributed training environments
Seniority
Mid-level, hands-on IC