CareerPlanGet AI match score →

Lead Member of Technical Staff, Inference Infrastructure

San Francisco🌐 Remote💼 Full-time🗓 2026-04-28 → 2026-07-31

Core

Designing and operating high-performance, scalable machine learning systems for deploying large language models via API endpoints.

Role type

Lead Member of Technical Staff, Inference Infrastructure

Builds

AI platform delivering large language models through API endpoints

Domain

Artificial Intelligence / Large Language Models / Cloud Infrastructure

Deliverable

production ML models

Required skills

Kubernetes architecture and management, distributed systems design, GPU/TPU accelerator optimization, multi-cloud infrastructure strategy, Linux environment management, cost management for compute/storage/network, high-performance server development (Golang/C++)

Preferred skills

Technical leadership across multiple teams, mentoring engineers, cross-functional collaboration

Technologies

Kubernetes, GCP, Azure, AWS, OCI, Linux, Golang, C++, GPUs, TPUs

Responsibilities

Drive architecture and strategy for deploying optimized NLP models to production, lead design of customized deployments for customers, mentor engineers to raise technical bar, establish patterns and practices across engineering teams, conduct senior-level technical reviews

Seniority

Senior, hands-on IC with technical leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗