Research, Post-Training Data
Core
Designing and executing data collection, synthesis, and evaluation strategies to steer large language models toward human preferences, reasoning, and helpfulness.
Role type
Post-training researcher (data-centric AI)
Builds
High-quality post-training datasets, evaluation pipelines, and metrics for model alignment
Domain
Artificial Intelligence / Large Language Models / Human-AI Interaction
Deliverable
production ML models
Required skills
Python, deep learning frameworks (PyTorch/TensorFlow/JAX), data curation, human feedback integration, synthetic data generation, experimental design, statistical analysis
Preferred skills
RLHF/RLAIF, preference modeling, reward learning, managing large-scale annotation workflows, active learning, model-assisted labeling
Technologies
PyTorch, TensorFlow, JAX
Responsibilities
Design data collection and synthesis strategies combining human and synthetic data; Develop pipelines for scalable human and model-assisted labeling; Research and model human preferences to improve model reasoning and truthfulness; Iterate on evaluation metrics and benchmarks; Scale existing methodologies and develop new ones; Publish research and share code/datasets
Seniority
Mid-to-Senior IC (PhD or equivalent industry experience preferred)