Research, Post-Training
Core
Developing and tuning post-training recipes, evaluations, and methodologies to make raw AI models safe, useful, and collaborative for humans.
Role type
Research, Post-Training
Builds
Post-training pipelines, evaluation frameworks, and research outputs for collaborative general intelligence models.
Domain
Artificial Intelligence / Machine Learning
Deliverable
production ML models
Required skills
Python, deep learning frameworks (PyTorch/TensorFlow/JAX), distributed training, debugging, statistical analysis, experimental design
Preferred skills
RLHF/RLAIF, preference modeling, reward learning, human data collection management, alignment research, PhD in CS/ML/Physics/Mathematics
Technologies
PyTorch, TensorFlow, JAX
Responsibilities
Iterate on post-training recipes (datasets, stages, hyperparameters); define and optimize evaluation metrics; debug training configurations and analyze results; scale methodologies and explore new dataset types; publish research and share code/datasets.
Seniority
Individual Contributor (Research/Engineering blend)