Staff MLOps Engineer
Required skills
Deep hands-on experience building MLOps platforms, including model registries, feature stores, and ML pipeline orchestration Working knowledge of model serving patterns: real-time inference, batch prediction, A/B deployment, and deployment strategies AWS infrastructure experience (ECS/EKS, S3, IAM, networking) and comfort operating in a Databricks ecosystem or equivalent lakehouse architecture
Experience with model monitoring
model evaluation, data drift detection, prediction drift, and performance degradation alerting A track record of building something from zero and bringing it to a state where others could operate and extend it Experience in a regulated industry (fintech, financial services, healthcare) where model governance is a compliance requirement
Preferred skills
Familiarity with model risk management frameworks and the ability to connect governance practices to regulatory expectations Experience working simultaneously with research-oriented ML teams and production-oriented engineering teams, and understanding how their needs diverge Infrastructure-as-code fluency (Terraform) Experience with ClickHouse or similar OLAP engines for low-latency ML feature serving Blockchain or crypto domain knowledge Contributions to open-source MLOps tooling
Technologies
AWS, Databricks ecosystem, model registry, model risk management framework, observability stack, feature store, ML pipeline orchestration, model serving, model monitoring, model evaluation, data drift detection, prediction drift, performance degradation alerting, ClickHouse or similar OLAP engines, Terraform
Responsibilities
Define the target-state MLOps architecture for Elliptic, covering model training pipelines, serving infrastructure, monitoring, feature management, and governance, and produce the architecture decision records that inform investment decisions Make and document build-vs-buy-vs-stop recommendations with clear cost modelling and trade-off analysis, evaluating vendors, open-source tools, and managed services against Elliptic's constraints (AWS-primary, Databricks ecosystem) Work with InfoSec to improve the existing model registry and model risk management framework, closing identified gaps in metadata, lineage, approval workflows, and drift/bias detection Build model training pipelines, CI/CD for ML, and serving infrastructure, working directly with a small group of infrastructure engineers to ship production-grade platform capabilities Instrument observability across the ML lifecycle: training metrics, serving latency and throughput, data quality, and prediction drift, integrating with Elliptic's existing observability stack Work directly with data scientists and ML engineers across all four consumer groups to onboard them onto the platform, writing documentation, runbooks, and reference architectures that lower the barrier to self-service
Seniority
Staff
Domain
Financial crime, fintech, financial services, compliance, model governance, MLOps, machine learning infrastructure, model risk management, model registry, observability, feature store, model serving, model monitoring, data drift detection, prediction drift, performance degradation alerting, model evaluation, model training pipelines, CI/CD for ML, batch prediction, real-time inference, A/B deployment, deployment strategies, AWS, Databricks ecosystem, ClickHouse or similar OLAP engines, Terraform