CareerPlanGet AI match score →

Senior Software Engineer, AI Inference Systems

Toronto, Ontario, Canada💼 Full-time🗓 2026-06-10 → 2026-07-19

Core

Architect and implement high-performance AI inference stacks, optimizing GPU kernels and compilers to serve large-scale models with extreme efficiency.

Role type

Senior IC machine-learning systems engineer (inference)

Builds

High-performance inference frameworks (vLLM), GPU kernels, compiler infrastructure, and benchmarking tools

Domain

AI/ML Systems, High-Performance Computing, GPU Architecture

Deliverable

production ML models

Required skills

Python, C/C++, CUDA, distributed systems, parallel programming, GPU memory hierarchy, container orchestration (Kubernetes/Docker), profiling/debugging

Preferred skills

Go, Rust, ML compilers (Triton, MLIR/LLVM), speculative decoding, disaggregation techniques

Technologies

vLLM, SGLang, PyTorch, Nsight Systems, Nsight Compute, Docker, Kubernetes, Slurm, NCCL

Responsibilities

Profile and optimize inference framework (vLLM) with parallelism and disaggregation techniques; Develop and benchmark GPU kernels using fusion and autotuning; Build high-level DSLs and compiler infrastructure; Architect scheduling for containerized large-scale inference on GPU clusters; Contribute to MLPerf Inference benchmarking suite; Conduct and publish original research on ML Systems

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗