Machine Learning Operations Engineer II
Core
Building and supporting a mature ML platform to empower ML engineers with state-of-the-art processes, tooling, and infrastructure for rapid iteration and reliable production of generative AI applications, agents, and data retrieval solutions.
Role type
MLOps Engineer II (Platform & Infrastructure)
Builds
Scalable, automated processes for model fine-tuning, reinforcement learning, LLM/Agent evaluation, and observability; internal tooling for the ML workflow.
Domain
Financial services / Generative AI / Machine Learning Operations
Deliverable
production ML models | infrastructure
Required skills
Kubernetes management, Python proficiency, distributed systems debugging, cloud platform (AWS) expertise, workflow orchestration (Ray, Airflow), open source evaluation and integration, LLM/Agent concepts.
Preferred skills
Experience with Agentic AI systems, MCP server patterns, running workflows on Ray.
Technologies
Python, Bash, LangGraph, PyTorch, Ray, Amazon EKS, Airflow, Jsonnet, Terraform, Git, Github, AWS, LangFuse, Sentry, Prometheus, W&B
Responsibilities
Iterate on ML processes to develop robust, auditable tools and frameworks; work closely with ML engineers to identify pain points and form effective solutions; empower engineers with stable tooling for rapid experimentation; provide resources and training on best practices for productionization; evaluate and champion open source/third-party solutions; ship scalable processes for model fine-tuning and RL; improve LLM and Agentic observability.
Seniority
Mid-level, hands-on IC