豆包AI大模型评测产品解决方案(Coding/Agent方向) - 火山方舟
Core
Design and execute evaluation solutions for enterprise clients in Coding/Agent scenarios to quantify LLM performance within their specific business workflows.
Role type
Senior Product Solutions Engineer (LLM Evaluation & Agent)
Builds
Quantifiable evaluation frameworks and case studies for enterprise AI adoption
Domain
Enterprise AI / Large Language Models / Software Development
Deliverable
production ML models
Required skills
Python/Go/Rust/Java, LLM evaluation methodologies, Agent tooling (Claude Code, Cursor, Codex, TRAE, DeepSeek Harness), benchmark analysis, root cause analysis, technical documentation
Preferred skills
Master's degree in CS/AI, experience with SWE-bench/HumanEval/NL2Repo, industry best practice formulation
Technologies
SWE-bench, HumanEval, Terminal Bench, NL2Repo, Cybergym, IDEs, CI/CD pipelines
Responsibilities
Guide clients in building high-quality evaluation schemes, design evaluation cases replicating end-to-end R&D workflows, collect feedback to drive model iteration, lead benchmark adoption for key clients, track and innovate on frontier evaluation technologies
