Infrastructure and MLOps Engineer
Core
Build, test, deployment, and productisation tools for Machine Learning Software components on HPC AI platforms.
Role type
Senior IC Infrastructure and MLOps Engineer
Builds
CI platform services, build engineering pipelines, component integration, and packaging/release systems for AI software.
Domain
AI Compute / High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Python, Linux environment management, CI/CD principles, Kubernetes, Docker, Cloud services (AWS), ML application maintenance, ML orchestration tools, ML accelerator hardware management
Preferred skills
Infrastructure as Code (Terraform/OpenTofu), GitHub Actions, Observability tooling (Prometheus), Grafana, Go/Java/C++
Technologies
Kubernetes, Docker, Terraform, AWS, Prometheus, Grafana, NV Ray, KFP, SkyPilot, DCGM
Responsibilities
Develop and maintain tools/services for AI research and engineering teams; Deploy and maintain services with Kubernetes and Docker; Manage Cloud Infrastructure using Terraform.
Seniority
Senior, hands-on IC