CareerPlanGet AI match score →

Software Engineer, ML Infrastructure

Toronto💼 Full-time🗓 2026-07-09 → 2026-07-31

Core

Build distributed training and inference infrastructure for generative AI models serving millions of users.

Role type

Software Engineer, ML Infrastructure

Builds

Distributed training systems and optimized inference pipelines for generative media models

Domain

Generative AI, Cloud Infrastructure, GPU/TPU Workloads

Deliverable

infrastructure

Required skills

Large-scale production infrastructure development, Distributed systems design, Kubernetes, GCP, GPU/TPU workload management, Linux environment troubleshooting, Worker scaling for training/inference, ML model runtime knowledge

Preferred skills

None stated

Technologies

Kubernetes, GCP, Linux

Responsibilities

Design distributed training infrastructure, Optimize inference pipelines, Deploy and troubleshoot complex Linux computing environments, Scale workers for training or inference workloads

Seniority

Mid-level, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗