AI Infrastructure Engineer, pAGI
Core
Build and operate distributed systems infrastructure for large-scale model training, evaluation, and inference to improve reliability, throughput, and resource efficiency.
Role type
Senior IC AI Systems Engineer (Infrastructure)
Builds
Shared grading services, inference platforms, compute schedulers, and self-service research tooling
Domain
AI Infrastructure / Distributed Systems / GPU Compute
Deliverable
production ML models | infrastructure
Required skills
distributed systems engineering, performance optimization, ML infrastructure, GPU performance tuning, system debugging, observability design
Preferred skills
experience with training stacks, automated validation pipelines, cross-team collaboration
Technologies
Kubernetes, Ray, custom distributed training frameworks, GPU clusters
Responsibilities
Design and deploy infrastructure solutions for training bottlenecks, develop automated capacity management and health monitoring, optimize compute scheduling to reduce idle time, diagnose end-to-end performance issues, build self-service tools for experiment launch and validation
Seniority
Senior, hands-on IC