Staff AI Engineer, Model Post-Training and Alignment
Core
Designing, executing, and optimizing post-training pipelines for large language models to improve performance, controllability, domain adaptation, and reasoning capabilities.
Role type
Staff AI Engineer, Model Post-Training and Alignment
Builds
Production-grade inference deployment and specialized small models from scratch
Domain
Cryptocurrency / Large Language Models
Deliverable
production ML models
Required skills
Large model post-training pipeline execution, Direct Preference Optimization (DPO), Generalized Reward Policy Optimization (GRPO), Reinforcement Learning from AI Feedback (RLAIF), Reward Model development, Domain-specific data strategy, Low-latency inference optimization
Preferred skills
Training specialized small models from scratch, Reinforcement Learning fundamentals, vLLM, SGLang
Technologies
vLLM, SGLang
Responsibilities
Lead and execute full post-training pipelines including supervised fine-tuning and reinforcement learning methods; Design and implement advanced training paradigms like DPO and GRPO; Develop domain-specific data recipes and augmentation pipelines; Conduct post-training of specialized small models from scratch; Build and refine Reward Models; Design RLAIF closed-loop systems; Optimize inference efficiency and deploy models using low-latency serving frameworks; Evaluate model performance using automated benchmarks and feedback loops; Collaborate with research and infrastructure teams to productionize workflows.