大模型评测工程师-Data(北京/上海)
Core
Build and maintain a large-scale evaluation system for SOTA large language models (including multimodal and Agent-based models) to ensure accurate, reproducible, and calibrated assessment metrics.
Role type
Senior IC Large Model Evaluation Engineer (Data & Infrastructure)
Builds
Internal evaluation environments, automated analysis tools, and scalable benchmarking pipelines for multimodal and Agent capabilities.
Domain
Artificial Intelligence / Large Language Models / Multimodal Systems
Deliverable
production ML models
Required skills
Machine learning and deep learning theory, Large model architecture and inference, PyTorch, Multimodal evaluation techniques, Agent evaluation methods, Automated data analysis, Algorithm implementation
Preferred skills
Open-source benchmark integration (e.g., BFCL, BrowseComp), Inference optimization (vLLM, SGLang, TensorRT-LLM), Quantization, Parallel inference, Agent RL environment construction
Technologies
PyTorch, vLLM, SGLang, TensorRT-LLM, Multimodal benchmarks, Agent-based benchmarks
Responsibilities
Deploy SOTA models to internal evaluation environments and run scaled tests across diverse datasets; Adapt and integrate mainstream open-source evaluation sets (text, multimodal, Agent) into the internal system; Conduct evaluation of Agent capabilities including multi-turn interactions, tool use, and sandbox environments; Design algorithms and tools for automated quantitative analysis of evaluation results and root cause tracing; Explore scalable environments and unbiased reward signals for Agent RL.
