Orchestration Workload Engineer - ACE - AI Factory
Core
Expert in workload orchestration owning and advancing scheduler tech stack across High-Performance Computing (HPC) and AI environments to ensure efficient scheduling, policy management, and resource optimization.
Role type
Senior IC Workload Orchestration Engineer (SLURM/HPC/AI)
Builds
Multi-node CPU and GPU environments, SLURM architecture, hybrid AI/HPC workloads, and global cross-functional orchestration standards.
Domain
Life sciences/pharmaceutical R&D, High-Performance Computing (HPC), AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
SLURM architecture and optimization, Kubernetes integration, containerization (Singularity/Apptainer), GPU scheduling (NVIDIA MIG), high-speed interconnects (InfiniBand, RoCE), Infrastructure-as-Code (Ansible, Terraform), observability and telemetry, multi-tenant cluster optimization, technical mentorship
Preferred skills
Experience in life sciences or pharmaceutical R&D, leadership in global cross-functional initiatives, strategic vision for HPC-cloud convergence
Technologies
SLURM, Kubernetes, Singularity, Apptainer, NVIDIA MIG, InfiniBand, RoCE, MPI, NCCL, Ansible, Terraform, SlurmDBD
Responsibilities
Architect and scale SLURM across heterogeneous HPC and AI environments; design and tune advanced SLURM configurations including custom plugins and topology-aware scheduling; integrate containerization standards and Kubernetes for hybrid workloads; lead global initiatives to establish orchestration standards and policies; act as a technical mentor and coach for junior and mid-level engineers; partner with Observability Engineers to establish deep telemetry dashboards.
Seniority
Senior, hands-on IC with mentorship responsibilities