大模型评测算法工程师(J100902)
Core
Build and maintain evaluation systems for LLM/VLM/Agent models in healthcare scenarios, covering medical Q&A, health literacy, diagnostic assistance, report interpretation, and medication consultation.
Role type
Senior IC large model evaluation engineer (healthcare)
Builds
Data-driven, reproducible, and scalable evaluation frameworks for medical AI products
Domain
Healthcare + Large Language Models (LLM) / Vision-Language Models (VLM) / Agents
Deliverable
production ML models
Required skills
LLM/VLM/Agent evaluation, benchmark construction, rubric design, automated evaluation, error analysis, statistical analysis, data engineering, prompt engineering, RAG evaluation, multi-modal understanding
Preferred skills
LLM-as-Judge, preference alignment, automated data generation, case mining, capability boundary analysis, model comparison
Technologies
LLM, VLM, Agent, RAG, Python, AI tools
Responsibilities
Design evaluation frameworks for core medical scenarios; implement automated evaluation capabilities including risk identification and regression analysis; conduct error attribution and capability diagnosis to guide model training and product strategy; track and apply frontier evaluation methods like Agent and multi-modal testing
Seniority
Senior, hands-on IC
