Machine Learning Engineer - AI & ML Evaluation Frameworks
Core
Architect and build large-scale evaluation frameworks to interrogate unimodal ML systems and multi-modal foundation models, leading deep-dive ML evaluations and failure analysis to ensure health features are mathematically sound, demographically equitable, and clinically safe.
Role type
Senior IC machine learning engineer (evaluation & safety)
Builds
Scalable evaluation infrastructure, synthetic data pipelines, automated frameworks, and data adaptors for sensor fusion
Domain
Digital health + AI safety & model interpretability
Deliverable
production ML models
Required skills
ML engineering, failure analysis, LLM/diffusion model evaluation, Python (production-grade), data pipeline construction, automated evaluation systems, bias detection, demographic equity measurement
Preferred skills
LLM/agentic system evaluation, synthetic data generation, prompt engineering, parallel data processing (Spark, Kubernetes, Airflow), privacy-preserving ML (Federated Learning), AI safety, model interpretability, adversarial testing
Technologies
Python, Spark, Kubernetes, Airflow, LLMs, diffusion models
Responsibilities
Design robust methodologies and scalable frameworks to assess performance, reliability, and safety of traditional ML and foundation models; Drive failure analysis and build instrumentation to detect clinical hallucinations, reasoning flaws, and edge cases; Expand LLM/diffusion-based data generation pipelines; Build data adaptors and visualizers to fuse asynchronous time-series signals; Develop generalizable tools and metrics to discover biases and measure demographic equity; Translate evaluation results into actionable engineering insights for researchers and clinical experts
Seniority
Senior, hands-on IC
