Director, Platform Engineering
Core
Ensure end-to-end reliability of the compute and cloud platform underpinning the company's AI initiatives, including incident management, GPU fleet operations, and DevOps pipelines.
Role type
Director, Platform Engineering (SRE focus)
Builds
Reliable AI infrastructure for molecular design, autonomous labs, supply chain, and professional services teams
Domain
Life sciences, diagnostics, biotechnology, AI/ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
Platform reliability ownership, incident management lifecycle, GPU/accelerator fleet management, cloud infrastructure (AWS/Azure/GCP), container orchestration (Kubernetes), Infrastructure as Code (Terraform/Pulumi), observability (metrics/logging/tracing), SLO definition, team leadership, CI/CD pipeline ownership
Preferred skills
LLM and agentic systems experience, scientific/research computing environments, compliance-constrained environments (SOC 2, GxP)
Technologies
Kubernetes, Docker, Terraform, Pulumi, AWS, Azure, GCP, GPU clusters
Responsibilities
Own availability, performance, and recovery for the compute and cloud platform; lead incident lifecycle and establish SLOs; design and operate GPU/accelerator fleets for large-scale training and inference; manage application support and operational layers; own CI/CD and release engineering; serve internal AI customer teams; define reliability strategy for agentic systems
Seniority
Director, organizational scale leadership