CareerPlanSign in

Member of Technical Staff, Model Efficiency

New York🌐 Remote💼 Full-time🗓 2025-11-07 → 2026-09-26

Core

Engineer focused on improving LLM inference efficiency, reducing latency, and increasing throughput in production environments.

Role type

Senior IC machine-learning systems engineer (inference optimization)

Builds

High-performance LLM inference systems for developers and enterprises

Domain

Artificial Intelligence / Large Language Models / Systems Engineering

Deliverable

production ML models

Required skills

C++ or Python, LLM inference ecosystem knowledge, performance bottleneck diagnosis, high-performance code development

Preferred skills

GPU programming (CUDA), low-level systems optimization, MoE architectures, speculative decoding, KV-cache optimization, distributed systems scaling

Technologies

vLLM, SGLang, CUDA, Transformers

Responsibilities

Dive deep into model execution to identify and resolve performance bottlenecks, collaborate with modeling and systems teams to experiment and ship inference improvements, develop innovative optimizations for GPU/CUDA and kernel-level performance

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.