Senior Machine Learning Operations Engineer
Core
Build and operate a low-latency, highly available real-time inference service for the risk decision engine.
Role type
Senior Machine Learning Operations Engineer
Builds
Production ML inference service, model deployment infrastructure, observability platform
Domain
Financial Risk / Machine Learning Operations
Deliverable
production ML models
Required skills
Python backend engineering, API frameworks (FastAPI/Flask), model registries, CI/CD pipelines, versioning, shadow mode, canary deployments, champion/challenger experiments, SHAP explainability, production observability, latency monitoring, error monitoring, drift detection, SQL, Redis, DynamoDB, Kafka, Kinesis, Redpanda
Preferred skills
Snowflake, dbt, Dagster, Airflow, regulated/compliance environments, Haskell, React, TypeScript
Responsibilities
Build and operate low-latency real-time inference service; Own model deployment infrastructure including registries and CI/CD; Build production observability for availability, latency, and drift; Partner with Risk Data Science on handoffs; Implement experimentation and explainability outputs; Help shape and build the ML platform team
