CareerPlanGet AI match score →

Senior ML Systems Engineer, Frameworks & Tooling

Toronto, Ontario, Canada🌐 Remote💼 Full-time🗓 2026-06-12 → 2026-07-26

Core

Design and maintain the core components of a training framework for large-scale LLM training, connecting research ideas to thousands of GPUs.

Role type

Senior ML Systems Engineer (Frameworks & Tooling)

Builds

Distributed training abstractions, monitoring/logging/debugging tooling, and robust systems for reproducible large-scale runs.

Domain

Large-scale AI model training, distributed systems, HPC infrastructure

Deliverable

production ML models

Required skills

Large-scale distributed training, HPC systems, JAX internals, multi-node cluster orchestration, CUDA/NCCL debugging, containerized environments, performance engineering

Preferred skills

LLM training experience, ML framework contributions (PyTorch, JAX, DeepSpeed, Megatron), evaluation/serving frameworks, data pipeline optimization

Technologies

JAX, Slurm, Ray, Kubernetes, Docker, Singularity/Apptainer, CUDA, NCCL, GB200, AMD, H200/100

Responsibilities

Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO), improve training throughput on multi-node clusters, develop tooling for monitoring and debugging, collaborate with infra teams on cluster/hardware configurations, resolve performance bottlenecks, build robust systems for reproducible runs

Seniority

Senior, hands-on IC

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗