CareerPlanSign in

实习自动驾驶感知3D算法工程师(J71300)

北京市💼 Full-time🗓 2026-08-20 → 2026-09-28

Core

Develop scene understanding algorithms based on Vision-Language Models (VLM) and multimodal large models for complex real-world scenarios, focusing on target, attribute, behavior, and semantic understanding.

Role type

IC machine-learning engineer (VLM/multimodal scene understanding)

Builds

VLM-based scene understanding models and data pipelines for autonomous driving perception

Domain

Autonomous driving, computer vision, multimodal AI

Deliverable

production ML models

Required skills

Python, C++, Linux, Transformer, ViT, LLM, VLM architectures, multimodal pretraining, visual-language alignment, SFT, RL, data mining, automatic labeling, synthetic data generation

Preferred skills

3D scene understanding, top-tier conference publications (CVPR, ICCV, NeurIPS, etc.), large-scale data loop experience

Technologies

VLM, LLM, Transformer, ViT, Python, C++, Linux

Responsibilities

Design and optimize VLM training and engineering deployment; explore fusion of VLM with traditional vision models; build automated data construction and data loop systems; track and validate frontier technologies in scene understanding

Seniority

Intern

Sourced via baidu · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.