Sr. Engineering Manager, MLOps
Core
Architecting and scaling the end-to-end infrastructure for production-grade Machine Learning, enabling seamless model development, training, deployment, and monitoring for Quince's Data Scientists and AI Researchers.
Role type
Senior Engineering Manager, MLOps
Builds
Automated, self-service MLOps platforms for model lifecycle management (training, deployment, serving) and high-throughput data pipelines.
Domain
Retail e-commerce, AI/ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
MLOps platform architecture, cloud-native infrastructure (AWS, Kubernetes, Docker), Infrastructure as Code (Terraform/Pulumi), ML frameworks (PyTorch, TensorFlow, Kubeflow, SageMaker), Feature Stores, high-throughput data pipelines (Spark, Flink, Kafka), CI/CD for ML, GPU cluster management, operational leadership
Preferred skills
LLM-ops, vector databases, real-time feature engineering
Technologies
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
Responsibilities
Define long-term MLOps roadmap transitioning workflows to automated self-service platforms; Build and maintain end-to-end infrastructure for model training, deployment, and serving; Partner with business leaders to align infrastructure investments with e-commerce drivers; Make high-judgment decisions on build vs. buy for cloud-native services; Oversee uptime and performance of production ML services; Direct optimization of high-cost computational resources (GPU clusters); Recruit and mentor ML Infra and DevOps engineers; Drive adoption of best practices in CI/CD, IaC, and automated testing; Translate AI Research needs into actionable engineering requirements; Establish KPIs for infrastructure team (deployment frequency, MTTR, cost per inference); Lead root-cause analyses for production failures; Monitor emerging trends in LLM-ops and real-time feature engineering.
Seniority
Senior, hands-on IC with management responsibilities