CareerPlanSign in

AI Engineer, Inference

Sydney, New South Wales, Australia💼 Full-time🗓 2026-09-29 → 2026-09-30

Core

Build and improve self-hosted AI inference services, establishing the engineering foundation for model serving in the organization's AI-factory environment to provide reliable, secure, scalable, and high-performance endpoints.

Role type

Senior AI Engineer (Inference)

Builds

Self-hosted model serving infrastructure, inference endpoints, and deployment templates for internal products and future Inference-as-a-service offerings.

Domain

AI Infrastructure / High-Performance Computing

Deliverable

production ML models

Required skills

Model serving optimization, distributed inference, Kubernetes, inference frameworks (TensorRT-LLM, vLLM, Triton), performance benchmarking, quantization, GPU memory management, observability

Preferred skills

Agentic workflows, speculative decoding, topology-aware placement

Technologies

TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, CUDA, cuDNN, NCCL, Kubernetes

Responsibilities

Build and operate self-hosted AI inference services; Define and implement standard model-onboarding workflows; Provision and manage secure, scalable inference endpoints; Develop reusable deployment templates, APIs, and SDKs; Optimize model-serving performance using quantization, compilation, and batching; Design distributed inference configurations; Work with Kubernetes and scheduler teams to define resource profiles; Build benchmarking and qualification workflows; Establish automated performance-regression testing; Build operational observability for inference services.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 891,000+ jobs from 20+ sources.