Machine Learning Research Scientist, Post-Training
Core
Develop novel post-training techniques (SFT, RLHF, reward modeling) to optimize data curation and evaluation for large-scale generative models in text and multimodal modalities.
Role type
Research Scientist (LLM Post-Training)
Builds
High-quality data and evaluation frameworks for next-generation generative AI models
Domain
Generative AI, Large Language Models, Reinforcement Learning
Deliverable
production ML models
Required skills
Deep learning, reinforcement learning, large-scale model fine-tuning, post-training techniques (RLHF, preference modeling, instruction tuning), bias mitigation, model robustness analysis
Preferred skills
Published research in top-tier AI conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR), customer-facing experience
Technologies
LLMs, SFT, RLHF, Reward Modeling
Responsibilities
Research and develop novel post-training techniques; Design and experiment new approaches to preference optimization; Analyze model behavior to identify weaknesses and propose solutions; Publish research findings in top-tier AI conferences
Seniority
Senior, hands-on IC