Senior/Staff Machine Learning Engineer - Health Evaluation - AI Teams (x/f/m)
Core
Designing, implementing, and scaling evaluation frameworks to ensure AI Health Companion systems behave safely, reliably, and helpfully for patients and practitioners.
Role type
Senior/Staff Machine Learning Engineer (LLM Evaluation & Agentic Systems)
Builds
Automated evaluation pipelines, metrics, protocols, and datasets for agentic AI systems
Domain
Healthcare AI, Large Language Models (LLMs), Agentic Systems
Deliverable
production ML models
Required skills
LLM evaluation, agentic system evaluation, experiment design, metric definition, evaluation automation, cross-functional collaboration
Preferred skills
Clinical or medical domain experience, sensitivity to ethical/regulatory challenges in healthcare AI
Technologies
Python, TypeScript, Java, Rails, React Native, LLMs (GPT, Claude, Llama, BERT)
Responsibilities
Define and own evaluation strategy (metrics, protocols, datasets, tooling); Implement and maintain automated evaluation pipelines; Run systematic experiments to assess reasoning, factuality, robustness, and UX; Collaborate with model developers and research scientists; Contribute to research on LLM evaluation methodologies
Seniority
Senior/Staff, hands-on IC with strategy influence