Member of Technical Staff, Model Efficiency
Core
Engineer focused on improving LLM inference efficiency, reducing latency, and increasing throughput in production environments.
Role type
Senior IC machine-learning systems engineer (inference optimization)
Builds
Optimized LLM inference stack components for enterprise customers
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
C++ or Python, LLM inference ecosystem knowledge, performance bottleneck diagnosis, high-performance code development
Preferred skills
GPU/CUDA programming, kernel-level optimization, MoE architectures, speculative decoding, KV-cache optimization, distributed systems scaling
Technologies
vLLM, SGLang, CUDA, Transformers
Responsibilities
Dive deep into model execution to identify and resolve performance bottlenecks, collaborate with modeling and systems teams to experiment and ship inference improvements, develop innovative optimizations for GPU/CUDA and kernel-level performance
Seniority
Senior, hands-on IC