Staff Engineer - ML Infra / MLOps
Core
Design and operate production-grade ML infrastructure platforms enabling distributed training, feature stores, and high-throughput inference serving for a scaling retail AI team.
Role type
Staff ML Infrastructure Engineer (MLOps)
Builds
End-to-end ML platform including training pipelines, serving infrastructure, and developer tooling for data scientists.
Domain
Retail / Cloud-native ML Infrastructure
Deliverable
production ML models
Required skills
Distributed training pipelines, Feature stores, High-throughput inference serving, GPU utilization optimization, Cloud cost control, Kubernetes (EKS), Infrastructure as Code (Terraform/Pulumi), ML frameworks (PyTorch, TensorFlow, Kubeflow, SageMaker), Data pipelines (Spark, Flink, Kafka), CI/CD for ML, Model versioning, Experiment tracking, Deployment strategies (blue-green, canary), Root-cause analysis, On-call discipline
Preferred skills
Startup hustle, Handling ambiguity, Rapid experimentation
Technologies
AWS, Kubernetes, Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
Responsibilities
Architect end-to-end ML platform design for scalability and extensibility; Build developer experience tooling to reduce friction from idea to production; Set engineering standards for CI/CD, IaC, and deployment strategies; Lead technical evaluation of core platform components; Optimize GPU utilization and cloud costs; Ensure production reliability and automated recovery; Mentor junior and mid-level engineers; Lead root-cause analyses for production failures.
Seniority
Staff, hands-on IC with mentorship
