Research Engineer/Scientist - Human Alignment, Consumer Devices
Core
Developing RLHF and post-training methods for personalized, multimodal AI systems to ensure long-term user alignment and beneficial behavior.
Role type
Research Engineer/Scientist (RLHF & Post-Training)
Builds
Adaptive, personalized AI models with long-term memory and user modeling capabilities.
Domain
Consumer Devices / Multimodal AI / Human Alignment
Deliverable
production ML models
Required skills
RLHF, reward modeling, preference optimization, post-training for large models, reinforcement learning, ranking, recommender systems, personalization, human-in-the-loop evaluation, dataset design, rubric creation, long-horizon evaluation, policy improvement, multimodal AI, training recipe development, data pipeline construction
Preferred skills
rigorous empirical work, clean experiment design, decision-useful metrics, nuanced behavioral objective training, cross-stack collaboration
Technologies
N/A
Responsibilities
Develop RLHF and post-training methods for multimodal models; Build reward models and preference-learning pipelines; Design datasets, rubrics, and evaluation frameworks; Run experiments on policy improvement using explicit and implicit feedback; Work on long-horizon evaluation problems; Collaborate with safety researchers to ensure alignment and constraints; Prototype and iterate on training recipes and evaluation suites; Define success metrics for personalized AI systems including trust and long-term benefit
Seniority
Senior, hands-on IC