CareerPlanSign in

Machine Learning Infrastructure Engineer, Model Inference

SF Office💼 Full-time🗓 2025-08-25 → 2026-09-26

Core

Design, deploy, and maintain scalable Kubernetes clusters and model serving infrastructure for real-time AI model inference and training in healthcare.

Role type

Senior IC ML Infrastructure Engineer (Model Inference)

Builds

Scalable Kubernetes clusters, high-performance model serving infrastructure, and robust model API orchestration systems.

Domain

Healthcare + Distributed Systems / Cloud Infrastructure

Deliverable

production ML models

Required skills

Kubernetes administration, distributed systems architecture, API development, GPU cluster management, compute-heavy workflow optimization

Preferred skills

NVIDIA Triton Server, VLLM, TRT-LLM, PyTorch, Tensorflow, CUDA optimization, Infrastructure as Code (Terraform, Ansible), GitOps, container registry management

Technologies

Kubernetes, NVIDIA Triton Server, VLLM, TRT-LLM, PyTorch, Tensorflow, Terraform, Ansible, CUDA

Responsibilities

Design and maintain scalable Kubernetes clusters for AI model inference and training; Develop and optimize ML model serving infrastructure for high-performance and low-latency; Collaborate with teams to scale backend infrastructure for model deployment and throughput optimization; Optimize compute-heavy workflows and enhance GPU utilization; Build a robust model API orchestration system; Define and implement strategies for scaling infrastructure as the company grows.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.