混元Agent评测Infra工程专家(北京/上海/深圳)
Core
Designing and engineering a platform to standardize and execute diverse Agent benchmarks (SWE, Terminal, Claw, MCP) in production environments.
Role type
Senior IC platform infrastructure engineer (Agent evaluation)
Builds
Scalable, reproducible evaluation platforms for AI agents
Domain
AI/LLM evaluation infrastructure
Deliverable
production ML models
Required skills
Backend system design, containerization and sandboxing, distributed task scheduling, concurrency control, network communication and proxy mechanisms, LLM/Agent benchmark logic abstraction
Preferred skills
Experience with SWE-bench, Terminal-Bench, MCP benchmarks
Technologies
Python, Go, Java, Kubernetes, Docker
Responsibilities
Integrate multiple Agent benchmarks into the evaluation platform; build the underlying runtime infrastructure (sandbox, dependency management, scheduling); ensure evaluation accuracy and observability; bridge algorithmic requirements with engineering implementation