Member of Technical Staff - Post Training, Applied (Vision)
Core
End-to-end ownership of vision-language model (VLM) post-training projects for enterprise customers, including data curation, fine-tuning, and evaluation.
Role type
Senior IC applied machine learning engineer (vision-language models)
Builds
Production-ready VLM adaptations and multimodal post-training pipelines for enterprise clients
Domain
Artificial Intelligence / Multimodal Machine Learning / Computer Vision
Deliverable
production ML models
Required skills
VLM post-training, supervised fine-tuning (SFT), preference alignment, reinforcement learning (RL), visual data curation, image-text pair curation, synthetic data generation, multimodal evaluation, visual grounding, OCR, document parsing, vision encoders, image-text architectures
Preferred skills
shared multimodal post-training infrastructure, customer-facing ML delivery, advanced alignment techniques
Technologies
SFT, RLHF, synthetic data generation pipelines, vision encoders, image-text architectures
Responsibilities
Translate customer requirements into multimodal post-training specifications, design and execute visual data generation and filtering processes, run supervised fine-tuning and alignment workflows, design task-specific evaluations for visual understanding and document parsing
Seniority
Senior, hands-on IC