Senior ML Infrastructure Engineer
Core
Build and operate high-performance GPU training and inference clusters to enable scientific breakthroughs in health, food security, climate, and AI.
Role type
Senior ML Infrastructure Engineer
Builds
Cloud and compute foundation for scientific research
Domain
Scientific research infrastructure / High-performance computing
Deliverable
infrastructure
Required skills
High-performance GPU cluster design and operation, High-throughput storage systems (Lustre), Distributed training optimization, GPU architecture expertise, High-speed networking, Infrastructure as Code (IaC), CI/CD practices
Preferred skills
Migration from traditional schedulers to containerized systems, Automated lifecycle management, Observability setup
Responsibilities
Build and optimize high-performance GPU clusters, Design high-throughput data paths, Benchmark and resolve performance bottlenecks, Establish observability and security controls, Partner with research teams for capacity forecasting
Seniority
Senior, hands-on IC