Staff Applied AI Scientist
Core
Building and maintaining production evaluation and observability frameworks for an AI Coach system to ensure continuous quality improvement and robust performance.
Role type
Staff Applied AI Scientist (LLMOps & Evaluation)
Builds
LLM-powered analysis tools, evaluation frameworks, and agentic orchestration systems for employee coaching.
Domain
HR Tech / People Science / Generative AI
Deliverable
production ML models | product features
Required skills
LLM evaluation (LLM-as-judge, human-in-the-loop), context engineering (RAG, memory, compression), agentic system design, observability tooling, longitudinal measurement, model selection and routing, guardrails and safety design, technical writing.
Preferred skills
Experience scaling eval practices across teams, public writing or talks in LLMOps, open-source contributions.
Technologies
Langfuse, Claude Code, Cursor, Codex, Python, LLMs.
Responsibilities
Own the end-to-end feedback loop for prompt engineering and evaluation; design and optimize context engineering for agentic flows; design and run longitudinal evaluation systems; contribute to agentic orchestration architecture; make model selection and routing decisions; create and monitor safety guardrails; enable other teams with reusable frameworks.
Seniority
Staff, hands-on IC with mentorship responsibilities.