Member of Technical Staff, Model Efficiency
Core
Engineer focused on improving LLM inference efficiency, reducing latency, and increasing throughput in production environments.
Role type
Senior IC machine-learning systems engineer (inference optimization)
Builds
High-performance LLM inference systems for developers and enterprises
Domain
Artificial Intelligence / Large Language Models / Systems Engineering
Deliverable
production ML models
Required skills
C++ or Python, LLM inference ecosystem knowledge, performance bottleneck diagnosis, high-performance code development
Preferred skills
GPU programming (CUDA), low-level systems optimization, MoE architectures, speculative decoding, KV-cache optimization, distributed systems scaling
Technologies
vLLM, SGLang, CUDA, Transformers
Responsibilities
Dive deep into model execution to identify and resolve performance bottlenecks, collaborate with modeling and systems teams to experiment and ship inference improvements, develop innovative optimizations for GPU/CUDA and kernel-level performance
Seniority
Senior, hands-on IC