Researcher, Post Training
Core
Designing new techniques for preference optimization, model evaluation, and feedback-driven learning to align multimodal models with human intent.
Role type
Researcher, Post-Training (Alignment & Evaluation)
Builds
Novel post-training methods, evaluation frameworks, and experimental systems for multimodal foundation models.
Domain
Generative AI, Multimodal Models, Model Alignment
Deliverable
production ML models
Required skills
Preference optimization (RLHF), model evaluation design, complex ML system debugging, training pipeline diagnostics
Preferred skills
Multimodal model training, open-source alignment contributions, human-in-the-loop evaluation systems
Technologies
RLHF, generative models, multimodal architectures
Responsibilities
Own research initiatives to improve model alignment and capabilities; Develop new post-training methods and evaluation frameworks; Partner with product and platform teams to define best practices; Implement and scale experimental systems for reliability; Translate research findings into production-ready systems.
Seniority
Individual Contributor (Researcher)