CareerPlanSign in

Machine Learning Operations Engineer II

Cambridge, Massachusetts; New York, New York💼 Full-time💰 $130–$175,000🗓 2026-04-16 → 2026-09-23

Core

Building and supporting a mature ML platform to empower ML engineers with state-of-the-art processes, tooling, and infrastructure for rapid iteration and reliable production deployment.

Role type

Senior MLOps Engineer (ML Platform & Infrastructure)

Builds

Internal tooling, services, and frameworks for the ML workflow; scalable processes for model fine-tuning, reinforcement learning, and LLM/Agent evaluation; observability solutions for agentic applications.

Domain

Financial services / Generative AI / Machine Learning Operations

Deliverable

production ML models | infrastructure | product features

Required skills

Kubernetes management, Cloud Platform (AWS), Python, distributed computing frameworks, workflow orchestration, software engineering best practices, debugging distributed systems, open source evaluation

Preferred skills

Agentic AI systems experience, Ray workflow experience, MCP server patterns, LLM/Agent concepts

Technologies

Python, Bash, LangGraph, PyTorch, Ray, Amazon EKS, Airflow, Jsonnet, Terraform, Git, Github, AWS, LangFuse, Sentry, Prometheus, W&B

Responsibilities

Iterate on ML processes to develop robust, auditable tools and services; work closely with ML engineers to identify pain points and form effective solutions; provide resources and training for ML teams on best practices; evaluate and champion open source and third-party solutions; ship scalable, automated processes for model fine-tuning and reinforcement learning; improve LLM and Agentic observability to monitor performance, decay, and drift.

Seniority

Mid-Senior, hands-on IC

Sourced via icims · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.