Applied Science PhD Intern - Clipchamp
Core
Research, prototype, and evaluate computer vision and multimodal models for video understanding, including agentic workflows for editing timelines.
Role type
PhD Research Intern (Computer Vision & Multimodal AI)
Builds
Agentic AI workflows, video understanding models, and scalable AI components for the Clipchamp product.
Domain
Computer Vision, Multimodal AI, Generative AI, Video Editing
Deliverable
production ML models | research
Required skills
Deep learning model training, computer vision, video understanding, agentic AI systems, tool use and function calling, multi-step orchestration, evaluation methodology, dataset curation, parameter-efficient fine-tuning
Preferred skills
Experience with transformers, attention mechanisms, vision-language models, diffusion models, transfer learning
Technologies
Python, PyTorch, Hugging Face Transformers, Diffusers, C#, TypeScript
Responsibilities
Design and build agentic workflows where language and multimodal models plan and act over an editing timeline; Fine-tune and benchmark state-of-the-art foundation models on domain data; Prepare and curate datasets for training and evaluation; Implement prototypes of scalable AI components; Share research findings through demos, write-ups, and presentations.
Seniority
Intern (PhD candidate)