Senior Research Engineer, LLM Training & Post-Training
Core
Design, build, and optimize training and post-training pipelines for large language models to improve model quality, training efficiency, and developer productivity.
Role type
Senior Research Engineer (LLM Training & Post-Training)
Builds
Production-ready AI systems, training infrastructure, and tooling for Lightning AI's platform and customer workloads.
Domain
Artificial Intelligence / Large Language Models / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Transformer-based LLM training and optimization, PyTorch, distributed training (multi-GPU), software engineering (Python), experiment design, model evaluation, debugging training issues
Preferred skills
SFT, RLHF, DPO, PPO, GRPO, DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, CUDA, Triton, vLLM, SGLang, TensorRT, GPU performance optimization
Technologies
PyTorch, DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, CUDA, Triton, vLLM, SGLang, TensorRT
Responsibilities
Design and optimize training/post-training pipelines; Improve model quality via fine-tuning and optimization; Build PyTorch-based training infrastructure; Optimize distributed training across multi-GPU environments; Investigate training convergence and performance bottlenecks; Design evaluation methodologies and benchmark models; Collaborate with customers to translate workloads into platform improvements; Partner with research and infrastructure teams to build production systems
Seniority
Senior, hands-on IC