CareerPlanGet AI match score →

Researcher, Post Training

*HQ - San Francisco, CA💼 Full-time🗓 2025-10-21 → 2026-07-31

Core

Designing new techniques for preference optimization, model evaluation, and feedback-driven learning to align multimodal models with human intent.

Role type

Researcher, Post-Training (Alignment & Evaluation)

Builds

Novel post-training methods, evaluation frameworks, and experimental systems for multimodal foundation models.

Domain

Generative AI, Multimodal Models, Model Alignment

Deliverable

production ML models

Required skills

Preference optimization (RLHF), model evaluation design, complex ML system debugging, training pipeline diagnostics

Preferred skills

Multimodal model training, open-source alignment contributions, human-in-the-loop evaluation systems

Technologies

RLHF, generative models, multimodal architectures

Responsibilities

Own research initiatives to improve model alignment and capabilities; Develop new post-training methods and evaluation frameworks; Partner with product and platform teams to define best practices; Implement and scale experimental systems for reliability; Translate research findings into production-ready systems.

Seniority

Individual Contributor (Researcher)

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗