CareerPlanGet AI match score →

Software Engineer, Inference AI/ML

Bellevue, WA💼 Full-time💰 $92,000–$92,000🗓 2026-07-01 → 2026-07-31

Core

Implement well-scoped features and fixes for model-serving services to improve latency, reliability, and cost on a GPU platform.

Role type

IC1 Software Engineer (Inference AI/ML)

Builds

Production features for model-serving services (Triton, vLLM, TensorRT-LLM, Ray Serve)

Domain

Cloud infrastructure for AI inference

Deliverable

production ML models

Required skills

Python, Go, C++, Linux fundamentals, Git/CI, data structures, algorithms, networked services

Preferred skills

PyTorch, TensorFlow, CUDA, Grafana, Prometheus, OpenTelemetry

Technologies

Triton, vLLM, TensorRT-LLM, Ray Serve, Kubernetes

Responsibilities

Implement features and fixes in Python/Go/C++ for model-serving services, Write tests, code comments, and short design docs, Add basic metrics and dashboards, Follow on-call runbooks and learn incident response, Contribute to performance experiments

Seniority

IC1, entry-level with mentorship

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗