CareerPlanSign in

Software Engineer - Model Performance

San Francisco💼 Full-time🗓 2024-03-28 → 2026-09-25

Core

Backend engineer optimizing ML model inference performance and infrastructure for LLMs.

Role type

Senior IC software engineer (ML performance)

Builds

High-performance inference stack for LLMs and embeddings

Domain

AI/ML infrastructure, Large Language Models

Deliverable

production ML models

Required skills

Python, C++, LLM optimization (quantization, speculative decoding, continuous batching), PyTorch, TensorRT, TensorRT-LLM, GPU architecture

Preferred skills

CUDA, Docker, Kubernetes, software engineering principles

Technologies

TensorRT, PyTorch, TensorRT-LLM, vLLM, sglang, CUDA

Responsibilities

Implement and productionize inference optimization techniques, debug ML performance issues in codebases, scale optimization across ML models, design and implement innovative solutions, own projects from idea to production

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.