大模型评测算法工程师-AI数据与安全
Core
Lead the construction and iteration of large model evaluation datasets, design automated analysis algorithms for defect localization, and develop red-blue teaming techniques to assess model security and robustness.
Role type
Senior IC large model evaluation algorithm engineer (AI safety & benchmarking)
Builds
Large model evaluation benchmarks, automated analysis tools, and security assessment frameworks
Domain
Artificial Intelligence / Large Language Models / AI Safety
Deliverable
production ML models
Required skills
Machine learning theory, deep learning fundamentals, large model architecture understanding, algorithm design, root cause analysis, adversarial testing design, technical research
Preferred skills
Large model benchmark construction, automated evaluation tool development, red-blue teaming experience, multi-modal model evaluation, top-tier conference publications
Technologies
PyTorch
Responsibilities
Design and iterate large model evaluation datasets with quality standards; Develop algorithms for automated evaluation result analysis and defect root cause tracing; Explore and implement adversarial testing and red-blue teaming to identify security vulnerabilities; Track and integrate frontier evaluation technologies and benchmark trends; Optimize evaluation tool algorithms to improve automation and efficiency