LLM Training Engineer
Required skills
Strong general software engineering skills (writing robust, performant systems), Experience with training or serving large neural networks (LLMs or similar), Solid grasp of deep learning fundamentals and modern literature, Comfort working in high-performance environments (GPU, distributed systems, etc.), Relevant experience (one or more) Pretraining / large-scale distributed training (FSDP/ZeRO/Megatron-style systems), Post-training pipelines (SFT, RLHF/RLAIF, preference optimization, eval loops), Building RL environments, simulators, or agent frameworks, Inference optimization, model compression, quantization, kernel-level profiling, Building large ETL pipelines for internet-scale data ingestion and cleaning, Owning end-to-end production ML systems with monitoring and reliability, Research orientation, Ability to propose and evaluate research ideas quickly, Strong experimental hygiene: ablations, metrics, reproducibility, analysis, Bias toward building — you can turn ideas into working code and results
Preferred skills
MS or PhD in Computer Science, Machine Learning, AI, Mathematics, or related field
Technologies
GPU, distributed systems, large-scale distributed training (FSDP/ZeRO/Megatron-style systems), post-training pipelines (SFT, RLHF/RLAIF, preference optimization, eval loops), RL environments, simulators, or agent frameworks, inference optimization, model compression, quantization, kernel-level profiling, large ETL pipelines for internet-scale data ingestion and cleaning, end-to-end production ML systems with monitoring and reliability
Responsibilities
Pretraining & Scaling, Post-training & RL, Sandbox Environments & Evaluation, Deployment & Inference Optimization
Seniority
Domain
AI infrastructure, multimodal AI models, serving platform