CareerPlanSign in

混元Agent评测Infra工程专家(北京/上海/深圳)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Designing and engineering a platform to standardize and execute diverse Agent benchmarks (SWE, Terminal, Claw, MCP) in production environments.

Role type

Senior IC platform infrastructure engineer (Agent evaluation)

Builds

Scalable, reproducible evaluation platforms for AI agents

Domain

AI/LLM evaluation infrastructure

Deliverable

production ML models

Required skills

Backend system design, containerization and sandboxing, distributed task scheduling, concurrency control, network communication and proxy mechanisms, LLM/Agent benchmark logic abstraction

Preferred skills

Experience with SWE-bench, Terminal-Bench, MCP benchmarks

Technologies

Python, Go, Java, Kubernetes, Docker

Responsibilities

Integrate multiple Agent benchmarks into the evaluation platform; build the underlying runtime infrastructure (sandbox, dependency management, scheduling); ensure evaluation accuracy and observability; bridge algorithmic requirements with engineering implementation

Sourced via tencent · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.