Sr MLOps Engineer
Core
Design, build, and maintain infrastructure and tools for the full machine learning lifecycle, ensuring seamless integration of ML models into production systems.
Role type
Senior MLOps Engineer
Builds
Production-grade ML infrastructure, orchestration tooling, and artifact/dataset storage solutions
Domain
Healthcare technology / Robotic surgery / Cloud infrastructure
Deliverable
infrastructure
Required skills
Kubernetes production operations, Python/Bash scripting, Infrastructure-as-Code (Terraform/Helm/Ansible), distributed storage systems, CI/CD pipelines, Linux system administration, GPU hardware management
Preferred skills
ML orchestration frameworks (Metaflow/MLflow/Kubeflow), large-scale infrastructure migration leadership, regulated industry experience
Technologies
Kubernetes, Metaflow, S3, MinIO, NVIDIA GPUs (B200/L40S/A6000/V100), CUDA, GitLab CI, ArgoCD
Responsibilities
Bootstrap and maintain production Kubernetes clusters with CNI and storage integration; Deploy and configure ML orchestration tooling and artifact storage; Validate GPU node health and configuration across heterogeneous hardware; Design and execute team migration playbooks for workflow and dataset porting; Write and maintain runbooks, architecture documentation, and disaster recovery procedures; Participate in on-call rotation and incident response; Collaborate with IT/Security on identity integration and compliance
Seniority
Senior, hands-on IC