CareerPlanSign in

大模型评测工程师-Data(北京/上海)

北京💼 Full-time🗓 2026-09-28

Core

Build and maintain a large-scale evaluation system for SOTA large language models (including multimodal and Agent-based models) to ensure accurate, reproducible, and calibrated assessment metrics.

Role type

Senior IC Large Model Evaluation Engineer (Data & Infrastructure)

Builds

Internal evaluation environments, automated analysis tools, and scalable benchmarking pipelines for multimodal and Agent capabilities.

Domain

Artificial Intelligence / Large Language Models / Multimodal Systems

Deliverable

production ML models

Required skills

Machine learning and deep learning theory, Large model architecture and inference, PyTorch, Multimodal evaluation techniques, Agent evaluation methods, Automated data analysis, Algorithm implementation

Preferred skills

Open-source benchmark integration (e.g., BFCL, BrowseComp), Inference optimization (vLLM, SGLang, TensorRT-LLM), Quantization, Parallel inference, Agent RL environment construction

Technologies

PyTorch, vLLM, SGLang, TensorRT-LLM, Multimodal benchmarks, Agent-based benchmarks

Responsibilities

Deploy SOTA models to internal evaluation environments and run scaled tests across diverse datasets; Adapt and integrate mainstream open-source evaluation sets (text, multimodal, Agent) into the internal system; Conduct evaluation of Agent capabilities including multi-turn interactions, tool use, and sandbox environments; Design algorithms and tools for automated quantitative analysis of evaluation results and root cause tracing; Explore scalable environments and unbiased reward signals for Agent RL.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.