CareerPlanGet AI match score →

Research Engineer/Scientist - Human Alignment, Consumer Devices

San Francisco💼 Full-time🗓 2026-03-11 → 2026-07-31

Core

Developing RLHF and post-training methods for personalized, multimodal AI systems to ensure long-term user alignment and beneficial behavior.

Role type

Research Engineer/Scientist (RLHF & Post-Training)

Builds

Adaptive, personalized AI models with long-term memory and user modeling capabilities.

Domain

Consumer Devices / Multimodal AI / Human Alignment

Deliverable

production ML models

Required skills

RLHF, reward modeling, preference optimization, post-training for large models, reinforcement learning, ranking, recommender systems, personalization, human-in-the-loop evaluation, dataset design, rubric creation, long-horizon evaluation, policy improvement, multimodal AI, training recipe development, data pipeline construction

Preferred skills

rigorous empirical work, clean experiment design, decision-useful metrics, nuanced behavioral objective training, cross-stack collaboration

Technologies

N/A

Responsibilities

Develop RLHF and post-training methods for multimodal models; Build reward models and preference-learning pipelines; Design datasets, rubrics, and evaluation frameworks; Run experiments on policy improvement using explicit and implicit feedback; Work on long-horizon evaluation problems; Collaborate with safety researchers to ensure alignment and constraints; Prototype and iterate on training recipes and evaluation suites; Define success metrics for personalized AI systems including trust and long-term benefit

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗