MTS, Research Engineer
Core
Designing novel AI model architectures and building distributed training infrastructure to enable state-of-the-art generative AI models.
Role type
Research Engineer (Generative AI Infrastructure)
Builds
High-performance distributed training systems and novel model architectures for LLMs
Domain
Generative AI, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
Python, C++, Rust, PyTorch, JAX, TensorFlow, CUDA, NCCL, MPI, linear algebra, calculus, probability, statistics, deep learning algorithm implementation
Preferred skills
low-level GPU programming (CUDA/Triton), hardware co-design, LLM training challenges, OSS inference engines (SGLang, vLLM)
Technologies
PyTorch, JAX, TensorFlow, CUDA, NCCL, MPI, Triton, SGLang, vLLM
Responsibilities
Explore new model architectures, training objectives, and optimization techniques; Implement and reproduce results from recent machine learning papers; Design, implement, and maintain high-performance, distributed machine learning systems; Translate abstract mathematical concepts into robust, efficient code; Collaborate with Research Scientists to unblock experiments and co-design hardware-aware experiments
Seniority
Mid-to-Senior, hands-on IC