Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Core
Design and productionize efficient LLM/VLM inference methods to solve frontier bottlenecks in compute, latency, and cost.
Role type
Senior Applied Scientist (Efficient LLM Inference & Model Optimization)
Builds
Production inference capabilities, quantized/distilled models, and optimized serving frameworks for Nebius Token Factory.
Domain
Cloud AI Infrastructure, Large Language Models (LLM), Vision Language Models (VLM), Model Optimization
Deliverable
production ML models
Required skills
LLM/VLM inference, transformer decoding algorithms, model compression, quantization, distillation, speculative decoding, KV-cache optimization, MoE routing, experimental design, statistical reasoning, Python, PyTorch
Preferred skills
vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, Triton, CUDA, PyTorch internals, post-training optimization (SFT/DPO/RLHF)
Technologies
PyTorch, Triton, CUDA, vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer
Responsibilities
Own research projects from hypothesis to production handoff; invent and evaluate inference optimization methods; build high-quality prototypes; design rigorous evaluation methodology covering latency, throughput, and cost; publish papers and technical reports; mentor engineers on experimental design.
Seniority
Senior, hands-on IC with research and production ownership