MTS, Research Engineer
Core
Design novel AI model architectures and build distributed training infrastructure to scale state-of-the-art research results.
Role type
Research Engineer (AI Infrastructure & Model Development)
Builds
Distributed training systems, high-performance ML frameworks, and scalable inference pipelines
Domain
Artificial Intelligence, Deep Learning, High-Performance Computing
Deliverable
production ML models
Required skills
Python, C++, Rust, PyTorch, JAX, TensorFlow, CUDA, NCCL, MPI, Linear Algebra, Calculus, Probability, Statistics
Preferred skills
Low-level GPU programming (Triton), LLM training challenges, OSS inference engines (SGLang, vLLM), Hardware co-design
Technologies
PyTorch, JAX, TensorFlow, CUDA, NCCL, MPI, Triton, SGLang, vLLM
Responsibilities
Design and implement novel model architectures and training objectives; Reproduce and extend state-of-the-art results from literature; Build and optimize distributed training systems for large GPU clusters; Translate research concepts into robust, efficient code; Collaborate with scientists to unblock experiments and co-design hardware-aware tools
Seniority
Mid-to-Senior, hands-on IC