实习自动驾驶感知3D算法工程师(J71300)
Core
Develop scene understanding algorithms based on Vision-Language Models (VLM) and multimodal large models for complex real-world scenarios, focusing on target, attribute, behavior, and semantic understanding.
Role type
IC machine-learning engineer (VLM/multimodal scene understanding)
Builds
VLM-based scene understanding models and data pipelines for autonomous driving perception
Domain
Autonomous driving, computer vision, multimodal AI
Deliverable
production ML models
Required skills
Python, C++, Linux, Transformer, ViT, LLM, VLM architectures, multimodal pretraining, visual-language alignment, SFT, RL, data mining, automatic labeling, synthetic data generation
Preferred skills
3D scene understanding, top-tier conference publications (CVPR, ICCV, NeurIPS, etc.), large-scale data loop experience
Technologies
VLM, LLM, Transformer, ViT, Python, C++, Linux
Responsibilities
Design and optimize VLM training and engineering deployment; explore fusion of VLM with traditional vision models; build automated data construction and data loop systems; track and validate frontier technologies in scene understanding
Seniority
Intern