Staff Software Engineer, Inference Infrastructure
Core
Designing, deploying, and operating high-performance, scalable machine learning systems for serving large language models via API endpoints.
Role type
Staff Software Engineer, Inference Infrastructure
Builds
AI platform delivering large language models through easy-to-use API endpoints
Domain
Artificial Intelligence / Large Language Models / Cloud Infrastructure
Deliverable
production ML models
Required skills
distributed systems design, Kubernetes orchestration, GPU workload management, multi-cloud infrastructure (GCP/Azure/AWS/OCI), Linux system administration, compute/storage/network resource optimization, high-performance server development (Golang/C++)
Preferred skills
familiarity with accelerator characteristics (GPUs/TPUs/custom accelerators), customer-facing deployment customization
Technologies
Kubernetes, GCP, Azure, AWS, OCI, Golang, C++, Linux
Responsibilities
Develop and deploy optimized NLP models to production environments, operate AI platform for low latency/high throughput, interface with customers for customized deployments, troubleshoot complex Linux-based computing environments, manage compute/storage/network resources and costs
Seniority
Staff, hands-on IC with strategic impact