Machine Learning Infrastructure Engineer
Core
Designing, building, and maintaining training and serving infrastructure for ML research.
Role type
ML Infrastructure Engineer
Builds
Training and serving infrastructure for ML research
Domain
Cloud infrastructure, High-performance computing, GPU clusters
Deliverable
infrastructure
Required skills
Cloud platform management, Kubernetes, GPU cluster management, Tooling development for diagnostics, Experiment management, GPU utilization optimization
Preferred skills
Large GPU cluster management, High-performance networking, LLM training support, ML framework development (PyTorch/TensorFlow/JAX), GPU kernel development
Technologies
Compute Engine, Kubernetes, Cloud Storage, PyTorch, TensorFlow, JAX
Responsibilities
Provide infrastructure support to ML research and product teams, Build tooling to diagnose cluster issues and hardware failures, Monitor deployments and manage experiments, Maximize GPU allocation and utilization for serving and training
Seniority
Senior, hands-on IC