Research Engineer II - ML Ops
Core
Build reliable AWS-based cloud infrastructure and MLOps pipelines to enable researchers to develop, test, and deploy cutting-edge ML solutions.
Role type
Senior IC Machine Learning Operations Engineer (Cloud Infrastructure)
Builds
Scalable AWS environments, MLOps pipelines, and internal tools for model training, evaluation, and deployment.
Domain
Cloud Infrastructure & Machine Learning Operations
Deliverable
production ML models
Required skills
AWS cloud services (EC2, ECS, EKS, Lambda, S3, EFS, DynamoDB, SageMaker), Infrastructure as Code (Terraform, Ansible, CloudFormation, CDK), Containerization (Docker, Kubernetes), CI/CD (GitHub Actions, GitLab CI, Jenkins), Linux administration, Software design patterns, API development (REST, gRPC), Monitoring (Prometheus, Datadog)
Preferred skills
Distributed computing frameworks (Spark, Ray), High-performance computing (HPC), Data governance, Open-source contributions
Technologies
AWS, Terraform, Ansible, CloudFormation, CDK, Docker, Kubernetes, GitHub Actions, GitLab CI, Jenkins, Prometheus, Datadog, SageMaker, Spark, Ray
Responsibilities
Migrate and optimize research workloads on AWS; Design and maintain MLOps pipelines for model lifecycle management; Develop prototypes and tools to accelerate research workflows; Translate research objectives into scalable engineering solutions; Mentor early-career engineers and promote engineering best practices.
Seniority
Mid-Senior, hands-on IC with mentorship responsibilities