Software Engineer 5 – Training Platform, AI Platform
Core
Design and build scalable infrastructure for large-scale machine learning model training, fine-tuning, and evaluation workflows.
Role type
Senior IC distributed ML infrastructure engineer
Builds
Training platform for foundation models and generative AI
Domain
Cloud computing, distributed systems, machine learning infrastructure
Deliverable
production ML models
Required skills
distributed model training, Kubernetes, Ray clusters, PyTorch, cloud computing (AWS), observability, logging, reporting, on-call processes, large-scale distributed training, parallelism techniques (FSDP, tensor/pipeline), GPU utilization optimization, memory efficiency, fault tolerance, checkpointing, cluster utilization
Preferred skills
modern ML model development workflows, cloud-based AI/ML services (SageMaker, Bedrock, Databricks, OpenAI), distributed training performance analysis tools (PyTorch Profiler, NVIDIA Nsight Systems), Generative AI training, fine-tuning, distillation
Technologies
Kubernetes, Ray, PyTorch, AWS, SageMaker, Bedrock, Databricks, OpenAI, PyTorch Profiler, NVIDIA Nsight Systems
Responsibilities
Design and build platform infrastructure, libraries, and SDKs for large-scale model training; Enable reliable and efficient training workflows for foundation models and generative AI; Diagnose and optimize performance of large distributed training jobs; Design easy-to-use APIs and interfaces for ML practitioners and non-experts
Seniority
Senior, hands-on IC