ML Operations Engineer
Core
Build and scale a high-performance machine learning and AI platform spanning on-prem data centers and GPU resources to support Data Science and AI teams.
Role type
Senior IC MLOps Engineer
Builds
Production ML/AI platform, model registries, container image pipelines, automated retraining triggers, and LLM/generative AI infrastructure (fine-tuning, vector databases, RAG).
Domain
Healthcare advertising, distributed systems, GPU infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Spark, Docker, Kubernetes, CI/CD, Linux administration, MLflow, Argo/Airflow, GPU orchestration, LLM/generative AI concepts
Preferred skills
CUDA fundamentals, NVIDIA driver/toolkit management, Jupyter/notebook workflows
Technologies
Python, Spark, Kubernetes, Docker, Argo Workflows, Airflow, MLflow, Prometheus, Grafana, NVIDIA CUDA, cuDNN
Responsibilities
Partner with Data Science/AI teams to adopt MLOps best practices and migrate workloads; Implement model tracking, versioning, and observability; Build CI/CD pipelines for ML artifacts; Improve platform reliability and cost-efficiency; Establish monitoring for GPU utilization and model metrics; Design and operate GPU cluster architecture and model serving infrastructure; Build infrastructure for LLM and generative AI workloads.
Seniority
Senior, hands-on IC