Machine Learning Operations Engineer II
Core
Building and supporting a mature ML platform to empower ML engineers with state-of-the-art processes, tooling, and infrastructure for rapid iteration and reliable production deployment.
Role type
Senior MLOps Engineer (ML Platform & Infrastructure)
Builds
Internal tooling, services, and frameworks for the ML workflow; scalable processes for model fine-tuning, reinforcement learning, and LLM/Agent evaluation; observability solutions for agentic applications.
Domain
Financial services / Generative AI / Machine Learning Operations
Deliverable
production ML models | infrastructure | product features
Required skills
Kubernetes management, Cloud Platform (AWS), Python, distributed computing frameworks, workflow orchestration, software engineering best practices, debugging distributed systems, open source evaluation
Preferred skills
Agentic AI systems experience, Ray workflow experience, MCP server patterns, LLM/Agent concepts
Technologies
Python, Bash, LangGraph, PyTorch, Ray, Amazon EKS, Airflow, Jsonnet, Terraform, Git, Github, AWS, LangFuse, Sentry, Prometheus, W&B
Responsibilities
Iterate on ML processes to develop robust, auditable tools and services; work closely with ML engineers to identify pain points and form effective solutions; provide resources and training for ML teams on best practices; evaluate and champion open source and third-party solutions; ship scalable, automated processes for model fine-tuning and reinforcement learning; improve LLM and Agentic observability to monitor performance, decay, and drift.
Seniority
Mid-Senior, hands-on IC
