Data Scientist
Core
Build a standardized evaluation layer to ensure safety, accuracy, and quality of AI conversations with patients across clinical, coaching, and support workstreams.
Role type
Junior Data Scientist (AI Safety & Evaluation)
Builds
Evaluation frameworks, monitoring services, and guardrail systems for patient-facing AI interactions.
Domain
Digital healthcare, AI safety, LLM evaluation
Deliverable
production ML models | product features
Required skills
Python, LLM fundamentals, prompt engineering, eval framework design, safety-first mindset, data modeling integration, metrics definition, automation
Preferred skills
Eval tooling (promptfoo, Ragas, LangSmith), healthcare domain experience, GCP, SQL
Responsibilities
Design and build a standardized evaluation framework for clinical and AI conversations; Build services to monitor patient-facing conversations and flag risky content; Develop eval datasets, metrics, and thresholds with clinicians and support leads; Integrate AI assistants with in-house patient data models; Operate systems in production with monitoring and alerting; Automate guardrail updates based on feedback.
Seniority
Junior, hands-on IC
