大模型平台研发(评测工程) - TikTok内容安全平台
Core
Building and maintaining a multi-modal large model evaluation system for TikTok's international content safety platform to detect risks and monitor ecosystem safety.
Role type
Senior IC machine-learning evaluation engineer (multi-modal content safety)
Builds
Automated evaluation frameworks, benchmarks, and diagnostic tools for text, image, and audio/video models
Domain
AI/ML, Content Safety, International Platforms
Deliverable
production ML models
Required skills
Python/Go/C++, LLM and multi-modal model principles, benchmark construction, error analysis, data cleaning and deduplication, automated evaluation framework development, root cause analysis
Preferred skills
LLM-as-Judge techniques, adversarial sample handling, long-tail risk assessment, Agent/LLM tool development
Technologies
Python, Go, C++, LLM frameworks, multi-modal training frameworks
Responsibilities
Design scientific and interpretable evaluation metrics and processes; Build high-quality automated evaluation sets and benchmarks; Develop automated evaluation frameworks for deployment and monitoring; Analyze evaluation results to identify model weaknesses and distribution shifts; Explore frontier evaluation technologies and convert them into reusable practices
Seniority
Senior, hands-on IC