CareerPlanSign in

Software Engineer (Model Inference)

Sydney, New South Wales💼 Full-time🗓 2026-09-17 → 2026-09-29

Core

Building and optimizing the end-to-end model inference stack to serve in-house and open-source AI models to tens of millions of users at low latency and high throughput.

Role type

Senior IC software engineer (ML inference & GPU systems)

Builds

High-throughput inference servers and optimized GPU serving frameworks

Domain

AI/ML inference, GPU-accelerated systems, interactive entertainment

Deliverable

production ML models

Required skills

GPU inference optimization, batching, quantization, CUDA kernel development, serving frameworks (vLLM, TensorRT, Triton), building software at scale

Preferred skills

Experience with LoRA testing and productionization, custom CUDA implementation

Technologies

vLLM, TensorRT, Triton, CUDA

Responsibilities

Design and ship high-throughput inference servers, optimize GPU utilization via batching and quantization, develop custom CUDA kernels, test and productionize models for millions of users

Seniority

Senior, hands-on IC

Sourced via viewjobs · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.