Member of Technical Staff - AI Evaluations, Health
Core
Own the end-to-end evaluation strategy for MAI Health's Mayo Clinic project, focusing on clinical accuracy, patient safety, and translating evaluation results into actionable product improvements for LLM-based medical tools.
Role type
Senior IC machine-learning engineer (health AI evaluations)
Builds
Layered evaluations spanning static benchmarks, adaptive multi-turn simulations, and agentic evaluations; scalable pipelines and dashboards for pre-launch and in-deployment performance.
Domain
Healthcare / Health Tech / Clinical AI
Deliverable
production ML models | dashboards & analysis
Required skills
LLM evaluation design, LLM-as-judge prompt engineering, user simulation, agentic/trajectory evaluation, software development, data science, machine learning, HIPAA/PHI handling, clinical workflow familiarity, medical terminology, stakeholder management, risk identification, cross-team coordination.
Preferred skills
Experience in regulated/safety-critical domains, familiarity with FDA guidance on AI-enabled devices, EU AI Act, lifecycle regulatory approaches, building ML/LLM-powered products.
Technologies
C, C++, C#, Java, JavaScript, Python
Responsibilities
Define metrics, build datasets, and translate results into actionable steps for response-quality improvement; work embedded with Mayo clinicians to design tooling and run evaluations; write LLM-as-judge prompts and build automated evaluations; evaluate patient-facing AI experiences for factuality, groundedness, and safety; serve as technical partner gathering requirements and triaging quality issues; partner with data teams to develop scalable pipelines; monitor trends in model performance and clinical-AI benchmarks.
Seniority
Senior, hands-on IC
