Senior MLOps Engineer
Core
Design and maintain infrastructure to move ML models from experimentation to reliable, scalable production deployment, including pipelines, serving, and observability.
Role type
Senior MLOps Engineer (Production AI Infrastructure)
Builds
Production ML pipelines, model serving platforms, and observability systems for AI applications
Domain
Data & AI, Cloud Infrastructure, Fintech, FMCG, Retail, Manufacturing
Deliverable
production ML models
Required skills
Python, ML lifecycle tools (MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML), model serving frameworks (BentoML, TorchServe, Triton), Docker, Kubernetes, CI/CD, Cloud platforms (AWS, Azure, GCP), Monitoring/Observability (Prometheus, Grafana)
Preferred skills
Infrastructure as Code (Terraform, Pulumi), Feature stores (Feast, Tecton), Data versioning (DVC, Delta Lake), Event-driven workflows (Kafka), LLM serving (vLLM, TGI), Model optimization (quantization, GPU tuning), RAG infrastructure, LLM evaluation tools
Technologies
MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, BentoML, TorchServe, Triton, Docker, Kubernetes, Prometheus, Grafana, Terraform, Pulumi, Feast, Tecton, DVC, Delta Lake, Kafka, vLLM, TGI, LangSmith, RAGAS
Responsibilities
Design and maintain infrastructure for ML model deployment; Create automated ML pipelines for training, evaluation, and retraining; Deploy and serve models across cloud, on-premise, or hybrid environments; Monitor model performance, data drift, and system health; Collaborate with Data Scientists and Engineers to improve model testing and maintenance; Establish best practices for CI/CD, model registry, and governance; Support GenAI and LLM-based systems at scale
Seniority
Senior, hands-on IC