Research Scientist - Audiovisual Understanding, Model Foundations
Core
Developing controllable, cutting-edge video generative models by merging algorithms with engineering to enhance machine learning components for training and inference.
Role type
Senior IC research scientist (large-scale video understanding & VLLMs)
Builds
State-of-the-art AI-generated video models and foundational models for LTX Studio
Domain
Generative AI, Computer Vision, Signal Processing, Large-scale Data Systems
Deliverable
production ML models
Required skills
Fine-tuning and controlling large-scale VLLMs, software engineering (Jax/PyTorch), computer vision model development, statistics, clustering, distributed systems implementation
Preferred skills
Post-training foundational models, data filtering and evaluation algorithms, system performance optimization
Technologies
Jax, PyTorch, VLLMs, distributed training frameworks
Responsibilities
Fine-tune and control VLLMs for video and audio understanding; Design algorithms for balancing, filtering, and curating training and evaluation datasets; Implement classic and modern algorithms for processing, clustering, evaluation and filtering of large scale datasets; Work within high-performance, scalable distributed systems capable of handling petabytes of data; Collaborate with researchers and product stakeholders to iteratively improve training sets and evaluation protocols