全栈研发工程师 - Seed Model
Core
Develop engineering infrastructure for multimodal generation model evaluation, including automated testing frameworks, result analysis, and platform product evolution.
Role type
Senior Full-Stack Engineer (LLM Evaluation & Agent Systems)
Builds
Automated evaluation systems, multimodal model testing platforms, and plugin-based delivery frameworks for AI models.
Domain
Artificial Intelligence / Large Language Models / Multimodal Systems
Deliverable
production ML models
Required skills
Full-stack development (frontend/backend), distributed system design, Agent framework development, LLM evaluation techniques, data structures and algorithms, storage and middleware technologies
Preferred skills
LLM training experience, LLM-as-a-Judge implementation, complex Harness framework construction, open-source community contributions
Technologies
Agent frameworks, LLM evaluation tools, distributed systems, frontend frameworks, storage systems, middleware
Responsibilities
Develop engineering infrastructure for multimodal model evaluation scenarios; Build sustainable and scalable automated evaluation systems based on Agent+Harness; Evolve evaluation platform capabilities towards Agent-driven user experiences; Construct plugin-based delivery systems with standardized modules and harness frameworks.
Seniority
Senior, hands-on IC
