CareerPlanGet AI match score →

Member of Technical Staff (AI Infrastructure Engineer)

London, England, UK💼 Full-time🗓 2026-06-10 → 2026-07-26

Required skills

Expert-level Kubernetes administration and YAML configuration management, Proficiency with Slurm job scheduling, resource management, and cluster configuration, Python and C++ programming with focus on systems and infrastructure automation, Hands-on experience with ML frameworks such as PyTorch in distributed training contexts, Strong understanding of networking, storage, and compute resource management for ML workloads, Experience developing APIs and managing distributed systems for both batch and real-time workloads, Solid debugging and monitoring skills with expertise in observability tools for containerized environments

Preferred skills

Experience with Kubernetes operators and custom controllers for ML workloads, Advanced Slurm administration including multi-cluster federation and advanced scheduling policies, Familiarity with GPU cluster management and CUDA optimization, Experience with other ML frameworks like TensorFlow or distributed training libraries, Background in HPC environments, parallel computing, and high-performance networking, Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices, Experience with container registries, image optimization, and multi-stage builds for ML workloads

Technologies

Kubernetes, Slurm, Python, C++, PyTorch, AWS

Responsibilities

Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads, Manage and optimize Slurm-based HPC environments for distributed training of large language models, Develop robust APIs and orchestration systems for both training pipelines and inference services, Implement resource scheduling and job management systems across heterogeneous compute environments, Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure, Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm, Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services, Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands

Seniority

Domain

AI, Machine Learning, Infrastructure, Cloud Computing

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗