Associate Director, Technical Lead (ML Engineering & Compute Platform Focus)
Core
Design, build, and operate ML engineering and compute platforms (on-prem HPC and cloud GPU) to enable scientific discovery at scale.
Role type
Senior IC ML Platform & Compute Engineer (Technical Lead)
Builds
Scalable ML training and inference infrastructure, MLOps pipelines, and hybrid compute environments.
Domain
Life Sciences / ML Infrastructure / Cloud Computing
Deliverable
infrastructure
Required skills
MLOps pipeline orchestration, GPU cluster management, AWS cloud compute (SageMaker, Batch, ParallelCluster), Kubernetes with GPU node management, containerization (Docker, NVIDIA Container Toolkit), Python, security implementation (OAuth2, OIDC, SAML, secrets management), AI-driven engineering tools
Preferred skills
NVIDIA DGX systems experience, distributed training frameworks, large-scale model training
Technologies
AWS, Kubernetes, Docker, NVIDIA Container Toolkit, SageMaker, Batch, ParallelCluster, Python
Responsibilities
Lead design and operation of hybrid ML compute platforms; implement MLOps pipelines; manage GPU resource allocation and job scheduling; embed security principles across the ML lifecycle; partner with data scientists to translate requirements into platform capabilities; champion AI-driven engineering transformation.
Seniority
Senior, hands-on IC with leadership responsibilities