Staff Research Engineer, Model Efficiency
Core
Develop, prototype, and deploy techniques to improve the speed and efficiency of Large Language Model (LLM) inference in production.
Role type
Staff Research Engineer (Model Efficiency)
Builds
Optimized LLM inference stack including model architecture, MoE routing, decoding algorithms, and software/hardware co-design.
Domain
Artificial Intelligence / Large Language Models / Model Efficiency
Deliverable
production ML models
Required skills
LLM architecture understanding, model inference optimization under resource constraints, model efficiency techniques, software engineering, technical mentorship
Preferred skills
PhD in Machine Learning, publications at top-tier conferences (ICLR, ACL, NeurIPS), experience in fast-paced startup environments
Technologies
GPU acceleration, MoE routing, decoding algorithms
Responsibilities
Optimize model architecture and MoE routing; improve decoding and inference-time algorithms; design software/hardware co-solutions for GPU acceleration; balance performance optimization with model quality.
Seniority
Staff, hands-on IC with mentorship