Software Dev Engineer II, Stores Foundational AI -SFAI
Core
Design and implement stable, efficient training systems and scalable data infrastructure for large-scale LLM training and reinforcement learning post-training.
Role type
Senior IC machine learning infrastructure engineer (RL post-training)
Builds
End-to-end RL post-training pipelines, training stability systems, and scalable data infrastructure for LLMs
Domain
Generative AI, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
Software development, System design and architecture, Machine Learning fundamentals, Transformer architecture, Reinforcement Learning, Data infrastructure design, GPU optimization, Observability system design
Preferred skills
ML frameworks (JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, TensorRT), CUDA/C++/Kernel development, System performance tuning, Parallel computing
Technologies
JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, TensorRT, CUDA, C++, PyTorch, RL algorithms (PPO, GRPO, RLOO)
Responsibilities
Design and build end-to-end RL post-training pipelines at cluster scale, Improve RL training stability by monitoring and tuning key metrics, Optimize RL post-training efficiency (GPU utilization, batching, sequence packing), Partner with research scientists to translate new RL algorithms into scalable systems, Profile and eliminate bottlenecks across compute, networking, and storage, Build observability systems for training dynamics and experiment tracking
Seniority
Senior, hands-on IC