多模态识别算法工程师-抖音直播
Core
Build multimodal content understanding, recognition, and mining models for live streaming to enhance content supply and growth.
Role type
Senior IC multimodal algorithm engineer (live streaming)
Builds
Production multimodal models for live streaming scenes
Domain
Live streaming, Computer Vision, NLP, Large Language Models
Deliverable
production ML models
Required skills
Computer vision, NLP, multimodal learning, deep learning, image/video understanding, object detection, segmentation, action recognition, RAG, few-shot learning
Preferred skills
PyTorch/TensorFlow, mixed precision training, distributed training, TensorRT deployment, video content understanding, multimodal retrieval, Kaggle/COCO/ActivityNet/ICPC/NOI/IOI awards
Technologies
PyTorch, TensorFlow, TensorRT
Responsibilities
Optimize and iterate computer vision, audio, and text large models for live streaming scenarios; design, develop, and tune algorithm models for cutting-edge technologies like CV, multimodal, and LLM.
Seniority
Senior, hands-on IC
