Applied Researcher, Vision Language Models/VLM - TikTok
Core
Researching and building state-of-the-art Vision Language Models (VLM) and foundation models to enhance user experiences in content moderation, search, and recommendations.
Role type
Senior IC Applied Researcher (Vision Language Models)
Builds
Multimodal reasoning and generation models, OCR/captioning features, and inference-efficient model designs for TikTok business applications.
Domain
Artificial Intelligence, Large Language Models, Multimodal AI
Deliverable
production ML models
Required skills
VLM pretraining and post-training, reinforcement learning alignment, model architecture design, distributed computing, deep learning frameworks (PyTorch, DeepSpeed, Megatron, vLLm), Python, Rust, C++
Preferred skills
Inference tuning and acceleration, GPU/AI accelerator expertise, PEFT, RL, MoE, CoT, Langchain, evaluation of AI systems, agent development
Technologies
PyTorch, DeepSpeed, Megatron, vLLm, PEFT, Langchain
Responsibilities
Lead research pushing state-of-the-art in multimodal reasoning and generation; Enhance VLMs with specialized features like OCR and captioning; Explore model architecture and inference-efficient designs; Collaborate with cross-functional teams to implement VLM projects; Extend research insights to academia.