Sr. Applied Scientist, Alexa AI
Core
Design, develop, and maintain end-to-end evaluation metrics and LLM-based evaluators to measure the quality of state-of-the-art conversational AI assistants.
Role type
Senior Applied Scientist (LLM Evaluation)
Builds
Production quality-evaluation metrics, LLM-as-a-Judge systems, and efficient evaluation models for conversational assistants.
Domain
Conversational AI, Large Language Models (LLMs), Natural Language Processing (NLP)
Deliverable
production ML models
Required skills
LLM architecture and training, Supervised Fine-Tuning (SFT), In-Context Learning (ICL), Learning from Human Feedback (LHF), model distillation, neural deep learning, GPU/TPU/Neuron deployment, data normalization and transformation
Preferred skills
vLLM, SGLang, TensorRT, conversational-assistant evaluation, production quality metric ownership
Technologies
Python, Java, C++, LLMs, GPUs, Neuron, TPU
Responsibilities
Define scientific direction for evaluation, mentor scientists and engineers, ensure data quality across acquisition and processing, partner with engineers on inference infrastructure, present data-backed proposals to leadership.
Seniority
Senior, hands-on IC with mentorship
