CareerPlanSign in

Staff Software Engineer, GPU Inference

Toronto Office💼 Full-time🗓 2026-07-28 → 2026-09-26

Core

Productionize and optimize GPU inference systems for large language models, focusing on disaggregated prefill/decode architectures and rack-scale AMD GPU infrastructure.

Role type

Staff Software Engineer (GPU Inference Systems)

Builds

GPU-accelerated inference stack (APIs, vLLM runtime, ROCm, rack-scale infrastructure)

Domain

AI Infrastructure / High-Performance Computing / GPU Systems

Deliverable

production ML models

Required skills

C++, Python, distributed systems debugging, GPU performance optimization, vLLM/SGLang/TensorRT-LLM, Linux/Kubernetes, numerical correctness validation, benchmarking

Preferred skills

AMD ROCm/HIP ecosystem, CUDA, open-source ML framework contributions, disaggregated prefill/decode architecture, KV-cache management, multi-GPU parallelism, quantization (BF16/FP8/INT8), kernel optimization

Technologies

vLLM, PyTorch, ROCm, HIP, Kubernetes, RDMA, BF16, FP8, INT8

Responsibilities

Design and deploy complete GPU prefill paths; establish operational readiness and automation for GPU fleets; define SLOs and improve fault isolation/recovery; profile and optimize throughput/latency/memory efficiency; tune scheduling and parallelism strategies; debug failures across application/runtime/hardware layers; build validation infrastructure for numerical correctness.

Seniority

Staff, hands-on IC with technical leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.