Lead Member of Technical Staff, Inference Infrastructure
Core
Designing and operating high-performance, scalable machine learning systems for deploying large language models via API endpoints.
Role type
Lead Member of Technical Staff, Inference Infrastructure
Builds
AI platform delivering large language models through API endpoints
Domain
Artificial Intelligence / Large Language Models / Cloud Infrastructure
Deliverable
production ML models
Required skills
Kubernetes architecture and management, distributed systems design, GPU/TPU accelerator optimization, multi-cloud infrastructure strategy, Linux environment management, cost management for compute/storage/network, high-performance server development (Golang/C++)
Preferred skills
Technical leadership across multiple teams, mentoring engineers, cross-functional collaboration
Technologies
Kubernetes, GCP, Azure, AWS, OCI, Linux, Golang, C++, GPUs, TPUs
Responsibilities
Drive architecture and strategy for deploying optimized NLP models to production, lead design of customized deployments for customers, mentor engineers to raise technical bar, establish patterns and practices across engineering teams, conduct senior-level technical reviews
Seniority
Senior, hands-on IC with technical leadership