CareerPlanSign in

大模型评测算法工程师(J100902)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Build and maintain evaluation systems for LLM/VLM/Agent models in healthcare scenarios, covering medical Q&A, health literacy, diagnostic assistance, report interpretation, and medication consultation.

Role type

Senior IC large model evaluation engineer (healthcare)

Builds

Data-driven, reproducible, and scalable evaluation frameworks for medical AI products

Domain

Healthcare + Large Language Models (LLM) / Vision-Language Models (VLM) / Agents

Deliverable

production ML models

Required skills

LLM/VLM/Agent evaluation, benchmark construction, rubric design, automated evaluation, error analysis, statistical analysis, data engineering, prompt engineering, RAG evaluation, multi-modal understanding

Preferred skills

LLM-as-Judge, preference alignment, automated data generation, case mining, capability boundary analysis, model comparison

Technologies

LLM, VLM, Agent, RAG, Python, AI tools

Responsibilities

Design evaluation frameworks for core medical scenarios; implement automated evaluation capabilities including risk identification and regression analysis; conduct error attribution and capability diagnosis to guide model training and product strategy; track and apply frontier evaluation methods like Agent and multi-modal testing

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.