Research Intern - AI/ML Numerics & Efficiency
Core
Research Intern role focused on advancing inquiry and theory in AI/ML numerics and efficiency through collaboration with doctoral candidates and researchers.
Role type
Research Intern (AI/ML Numerics & Efficiency)
Builds
Research and development strides in ML systems and model efficiency
Domain
Artificial Intelligence / Machine Learning Systems / High-Performance Computing
Deliverable
research
Required skills
Python, C++, machine learning systems, transformer-based model architectures, attention mechanisms, KV cache behavior, training and inference bottlenecks, PyTorch, Hugging Face Transformers, SGLang, vLLM, TensorRT-LLM, GPU or accelerator programming, CUDA, Triton, profiling and performance analysis, benchmarking, low-precision numeric, quantization methods, hardware-software co-design, ML systems, model optimization, kernel development, numerical computing
Preferred skills
open-source ML framework contributions, modern ML frameworks and runtimes, GPU or accelerator programming, profiling and performance analysis, benchmarking tools, low-precision numeric, quantization methods, hardware-software co-design, ML systems, model optimization, kernel development, numerical computing
Technologies
PyTorch, Hugging Face Transformers, SGLang, vLLM, TensorRT-LLM, CUDA, Triton
Responsibilities
learn, collaborate, and network with fellow doctoral candidates and researchers; present findings; contribute to the vibrant life of the community
Seniority
Intern