CareerPlanGet AI match score →

Staff Research Engineer, Model Efficiency

New York🌐 Remote💼 Full-time🗓 2025-11-07 → 2026-08-01

Core

Develop, prototype, and deploy techniques to improve the speed and efficiency of Large Language Model (LLM) inference in production.

Role type

Staff Research Engineer (Model Efficiency)

Builds

Optimized LLM inference stack including model architecture, MoE routing, decoding algorithms, and software/hardware co-design.

Domain

Artificial Intelligence / Large Language Models / Model Efficiency

Deliverable

production ML models

Required skills

LLM architecture understanding, model inference optimization under resource constraints, model efficiency techniques, software engineering, technical mentorship

Preferred skills

PhD in Machine Learning, publications at top-tier conferences (ICLR, ACL, NeurIPS), experience in fast-paced startup environments

Technologies

GPU acceleration, MoE routing, decoding algorithms

Responsibilities

Optimize model architecture and MoE routing; improve decoding and inference-time algorithms; design software/hardware co-solutions for GPU acceleration; balance performance optimization with model quality.

Seniority

Staff, hands-on IC with mentorship

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗