CareerPlanGet AI match score →

Senior Software Engineer, AI Inference

Toronto, Ontario, Canada💼 Full-time🗓 2026-06-14 → 2026-07-30

Core

Senior Software Engineer specializing in large-scale LLM inference systems, profiling deployments, and optimizing performance for NVIDIA's inference stack and open-source ecosystem.

Role type

Senior IC systems engineer (LLM inference)

Builds

High-performance LLM serving infrastructure, benchmarking campaigns, and internal tooling for GPU clusters.

Domain

AI/ML Infrastructure, Large Language Models, High-Performance Computing

Deliverable

production ML models

Required skills

LLM inference optimization, vLLM deployment, Kubernetes orchestration, Slurm scheduling, GPU performance profiling, system architecture at scale

Preferred skills

Kernel engineering collaboration, open-source contribution, performance modeling

Technologies

vLLM, Kubernetes, Slurm, Nsight Systems, Nsight Compute, GPU clusters

Responsibilities

Design and implement end-to-end benchmarking campaigns across Kubernetes and Slurm environments; Set up and operate vLLM serving deployments on GPU clusters and tune configurations for throughput and latency; Develop detailed performance plans based on profiling findings; Build internal tools, benchmarking harnesses, and automation pipelines; Document architectures, findings, and recommendations for technical audiences.

Seniority

Senior, hands-on IC

Rewrite
## About the Role Help us push the boundaries of AI inference at NVIDIA — where your systems expertise shapes both the technology and the teams building on top of it! We're looking for a Senior Software Engineer to work at the frontier of large-scale LLM serving, partnering directly with some of the world's most technically demanding customers to unlock the full performance potential of NVIDIA's inference stack. In this role, you'll combine deep systems knowledge with hands-on customer engagement — profiling real deployments, benchmarking across GPU clusters, and turning insights into improvements that ripple across the open-source ecosystem. Do you love digging into performance problems that don't have obvious answers, and want your work to have an impact far beyond a single codebase? We'd love to talk. Unlike traditional customer-facing engineering roles, we expect you to go far deeper — contributing to vLLM, NVIDIA Dynamo, and the tooling that makes every engineer on your team more effective. ## What You'll Be Doing - Work directly with customer engineering teams through long-term technical partnerships, understanding their LLM serving architectures and performance goals, then designing and implementing end-to-end benchmarking campaigns across Kubernetes and Slurm environments to surface actionable insights. - Set up and operate vLLM serving deployments on GPU clusters, tuning configurations for throughput, latency, and efficiency — and collect Nsight Systems / Nsight Compute profiling traces to identify performance gaps relative to reference frameworks. - Develop detailed performance plans based on profiling findings and collaborate with NVIDIA's kernel engineering and OSS vLLM teams to drive improvements that benefit both your customers and the broader community. - Build internal tools, benchmarking harnesses, and automation pipelines that raise the productivity of your teammates and customers alike — with a multiplier attitude that makes everyone around you more effective. - Document architectures, findings, and recommendations with clarity for technical audiences, and contribute improvements back to vLLM and related open-source projects where appropriate. ## What We Need To See - Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or equivalent experience. - 5+ years of industry experience building and operating complex, production-grade software systems, with strong instincts for how systems behave at scale. - Hands-on experience deploying and operating LLM inference workloads — particularly with vLLM — including configuration, optimization, and debugging in real-world environments. - Proficiency with container orchestration (Kubernetes) and HPC scheduling (Slurm) for running GPU-accelerated workloads. - Solid understanding of LLM serving fundamentals: batching strategies (continuous batching, chunked prefill), KV cache management, and tensor/pipeline parallelism. - Familiarity with GPU performance analysis: memory hierarchy, utilization, roofline modeling, and pr ## Nice to Have - Experience with open-source contributions and community engagement. - Strong communication skills to collaborate effectively with cross-functional teams and customers. - Passion for AI and machine learning technologies. - Experience with cloud platforms and distributed systems. ## Benefits - Opportunity to work on cutting-edge AI inference technologies. - Collaborate with top-tier customers and industry leaders. - Contribute to open-source projects and influence the broader AI community. - Access to NVIDIA's advanced tools and technologies. - Professional growth and development opportunities.
Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗