ML Ops Engineer
Core
Design, scale, and maintain high-performance MLOps and LLMOps infrastructure for an AI-infused scenario planning platform.
Role type
Senior IC ML Ops Engineer (LLMOps & Cloud Infrastructure)
Builds
Cloud-native AI/ML infrastructure, CI/CD pipelines, and production deployments for LLMs and generative AI workloads.
Domain
Enterprise SaaS, AI/ML Infrastructure, Cloud Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Docker, Terraform, Python, GPU orchestration, CI/CD pipelines, cloud cost optimization, observability
Preferred skills
vLLM, Ray, MLflow, LangChain, DeepSpeed, Hugging Face TGI, service meshes (Istio)
Technologies
Kubernetes, Docker, Helm, Terraform, Ansible, GitHub Actions, ArgoCD, Jenkins, Prometheus, Grafana, OpenTelemetry, MLflow, vLLM, Ray, Triton Inference Server, NVIDIA GPU Operator, Slurm, Ray, Kubecost
Responsibilities
Provision and manage cloud-native AI/ML infrastructure using Kubernetes and GPU orchestration frameworks; Automate core platform infrastructure using IaC tools; Optimize GPU compute workloads and high-speed networking; Build and maintain robust CI/CD and MLOps pipelines; Deploy LLMs and generative AI workloads using advanced inference engines; Enable automated model validation and monitoring for drift; Monitor and optimize cloud spend across GPU/CPU clusters; Implement auto-scaling strategies and dynamic resource allocation; Establish benchmarking and telemetry for model unit economics; Implement end-to-end observability.
Seniority
Senior, hands-on IC