CareerPlanGet AI match score →

Member of Technical Staff, Model Efficiency

Toronto, Ontario, Canada🌐 Remote💼 Full-time🗓 2026-06-11 → 2026-07-27

Core

Engineer focused on improving LLM inference efficiency, reducing latency, and increasing throughput in production environments.

Role type

Senior IC machine-learning systems engineer (inference optimization)

Builds

Optimized LLM inference stack components for enterprise customers

Domain

Artificial Intelligence / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

C++ or Python, LLM inference ecosystem knowledge, performance bottleneck diagnosis, high-performance code development

Preferred skills

GPU/CUDA programming, kernel-level optimization, MoE architectures, speculative decoding, KV-cache optimization, distributed systems scaling

Technologies

vLLM, SGLang, CUDA, Transformers

Responsibilities

Dive deep into model execution to identify and resolve performance bottlenecks, collaborate with modeling and systems teams to experiment and ship inference improvements, develop innovative optimizations for GPU/CUDA and kernel-level performance

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗