Staff Engineer - ML Operations - USA Remote
Core
Design and operate the ML lifecycle pipelines, serving infrastructure, and observability tools that enable scientists to run large-scale protein design and structure prediction experiments with speed and reproducibility.
Role type
Staff Engineer, MLOps (Infrastructure & Platform)
Builds
Scalable, reproducible ML pipelines for protein design and structure prediction; self-service ML tooling for data scientists.
Domain
Life Sciences / Bioinformatics / Computational Biology
Deliverable
production ML models | infrastructure
Required skills
MLOps lifecycle management, containerization and orchestration (Docker, Kubernetes), cloud platform expertise (Azure), Python, CI/CD automation, GPU workload management, observability and monitoring.
Preferred skills
Computational biology pipeline support, regulated environment experience (GxP, SOX, HIPAA), LLMOps and agentic frameworks.
Technologies
MLflow, Weights & Biases, Kubeflow, Airflow, Dagster, Prefect, Nextflow, OpenTelemetry, Prometheus, Grafana, Azure ML, Langfuse.
Responsibilities
Own end-to-end ML lifecycle including experiment tracking, model registry, versioning, and lineage; design and operate model serving for batch and low-latency online inference; implement CI/CD, continuous training, and observability; drive GPU and accelerated-compute efficiency; build self-service ML tooling and set MLOps standards.
Seniority
Staff, hands-on IC with technical leadership