Member of Technical Staff, North Modelling (Evals)
Core
Build evaluation systems, feedback loops, and applied modelling workflows to ensure model progress translates into better product outcomes for the North AI workspace platform.
Role type
Applied Machine Learning Engineer (Evaluation Strategy)
Builds
Evaluation systems, feedback loops, and applied modelling workflows for the North platform
Domain
Enterprise AI / Large Language Models / AI Agents
Deliverable
production ML models
Required skills
LLM evaluation strategy, feedback loop design, data curation, model selection, failure analysis, rubric creation, regression tracking, product-outcome reasoning
Preferred skills
Agent workflow evaluation, privacy-preserving data usage, qualitative signal translation, self-directed problem solving
Technologies
LLMs, AI agents, production modelling pipelines
Responsibilities
Define evaluation strategy for agent workflows and human-AI interactions; Build high-quality evals from user feedback and production data; Create systems to continuously update evals with product learning; Translate eval results into actionable recommendations for central modelling teams
Seniority
Mid-Senior, hands-on IC

