Research Scientist (Trust and Safety - Vision Language Models/VLM)
Core
Researching and enhancing Vision Language Models (VLMs) with specialized features like OCR and captioning to optimize performance for TikTok business applications including content moderation, search, and recommendations.
Role type
Research Scientist (Trust and Safety - Vision Language Models/VLM)
Builds
Foundation models (LLM, VLM, Omni Models) and downstream business applications for TikTok
Domain
Artificial Intelligence, Large Language Models, Vision Language Models, Trust and Safety
Deliverable
production ML models
Required skills
VLM pretraining and applications, multi-modality LLM/VLM expertise, reinforcement learning based alignment, efficient training and inference, model architecture design, OCR and captioning implementation, cross-functional project planning
Preferred skills
Inference tuning and acceleration, GPU/AI accelerator expertise, distributed computing framework tuning, PEFT, RL, MoE, CoT, Langchain, published research papers
Technologies
Python, Rust, C++, PyTorch, DeepSpeed, Megatron, vLLM, PEFT, Langchain
Responsibilities
Enhance VLM with specialized features like OCR and captioning, Explore model architecture and inference-efficient design, Work with cross-functional teams to implement VLM projects, Extend insights from industry to academia
Seniority
PhD level Research Scientist