CareerPlanSign in

Applied Researcher, Vision Language Models/VLM - TikTok

San Jose, United States of America💼 Full-time🗓 2026-09-28

Core

Researching and building state-of-the-art Vision Language Models (VLM) and foundation models to enhance user experiences in content moderation, search, and recommendations.

Role type

Senior IC Applied Researcher (Vision Language Models)

Builds

Multimodal reasoning and generation models, OCR/captioning features, and inference-efficient model designs for TikTok business applications.

Domain

Artificial Intelligence, Large Language Models, Multimodal AI

Deliverable

production ML models

Required skills

VLM pretraining and post-training, reinforcement learning alignment, model architecture design, distributed computing, deep learning frameworks (PyTorch, DeepSpeed, Megatron, vLLm), Python, Rust, C++

Preferred skills

Inference tuning and acceleration, GPU/AI accelerator expertise, PEFT, RL, MoE, CoT, Langchain, evaluation of AI systems, agent development

Technologies

PyTorch, DeepSpeed, Megatron, vLLm, PEFT, Langchain

Responsibilities

Lead research pushing state-of-the-art in multimodal reasoning and generation; Enhance VLMs with specialized features like OCR and captioning; Explore model architecture and inference-efficient designs; Collaborate with cross-functional teams to implement VLM projects; Extend research insights to academia.

Sourced via tiktok · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.