LLM Engineer (Reinforcement Learning)
Core
Designing LLM training pipelines to train generative language models for real-world service deployment and continuous quality improvement.
Role type
Senior IC LLM Engineer (Reinforcement Learning)
Builds
Production-ready generative language models and training infrastructure
Domain
Artificial Intelligence / Large Language Models / Reinforcement Learning
Deliverable
production ML models
Required skills
Deep Learning, NLP, Python, PyTorch, Distributed Training (Slurm, DDP, Horovod), GPU Cluster Management, Reward Hacking prevention, Self-Refine architecture design, Tool selection for LLMs
Preferred skills
Academic publications (ACL, EMNLP, NeurIPS), PhD, Docker, Kubernetes, Post-training experience, Supervised Fine-Tuning, Parameter Efficient Fine-Tuning
Technologies
PyTorch, Slurm, DDP, Horovod, Docker, Kubernetes, PPO, GRPO, DPO
Responsibilities
Optimize LLM training process efficiency, Improve accuracy and stability of generated results, Design learning structures to prevent reward hacking and enable self-refinement, Develop base models capable of integrating external knowledge and APIs, Train LLMs to autonomously select external tools based on instructions
Seniority
Senior, hands-on IC