CareerPlanSign in

Machine Learning Engineer

Canada💼 Full-time🗓 2026-06-04 → 2026-08-06

Core

Develop SOTA deep learning software and model optimization algorithms for LLM inference, focusing on compression, quantization, and speculative decoding to accelerate AI for enterprises.

Role type

Senior IC machine learning engineer (LLM inference optimization)

Builds

Production-ready LLM inference systems, model compression pipelines, and speculative decoding frameworks for enterprise deployment

Domain

Artificial Intelligence / Large Language Models / Model Optimization

Deliverable

production ML models

Required skills

Deep learning fundamentals, LLM inference optimization, tensor math libraries (PyTorch, NumPy), Python, algorithm design, linear algebra, mathematical modeling

Preferred skills

Experience with model quantization and pruning, speculative decoding frameworks, hardware profiling (CPU/GPU), open-source contribution

Technologies

PyTorch, NumPy, vLLM, LLM-compressor

Responsibilities

Design and implement model compression pipelines using quantization and pruning; Develop and maintain speculative decoding frameworks to improve inference speed; Profile and optimize end-to-end LLM performance including memory, latency, and throughput; Collaborate with research scientists to translate experimental ideas into production systems; Contribute to open-source projects and code reviews; Mentor team members

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.