大模型评测研发工程师-AI数据与安全
Core
Develop engineering infrastructure for large model evaluation, including dataset management, sampling, human/machine evaluation capabilities, and analysis; build AI Agents to automate evaluation workflows.
Role type
Senior IC large model evaluation engineer (AI infrastructure & Agents)
Builds
Evaluation infrastructure, automated evaluation Agents, and evaluation datasets
Domain
Artificial Intelligence / Large Language Models / AI Safety
Deliverable
production ML models
Required skills
Full-stack development, distributed system design, LLM principles, Agent framework development, evaluation methodology design, data pipeline engineering
Preferred skills
LLM training experience, LLM-as-a-judge implementation, multi-agent system design, open-source contributions
Technologies
Distributed systems, storage middleware, frontend frameworks, Agent frameworks
Responsibilities
Design and develop evaluation infrastructure for dataset ingestion and management; Build AI Agents to enable end-to-end automated evaluation; Analyze evaluation results and optimize evaluation metrics; Collaborate with algorithm and product teams to define evaluation standards.
Seniority
Senior, hands-on IC