Staff Software Engineer, Inference Infrastructure
Core
Building high-performance, scalable, and reliable machine learning systems to deploy optimized NLP models to production in low latency, high throughput, and high availability environments.
Role type
Staff Software Engineer, Inference Infrastructure
Builds
AI platform delivering large language models through API endpoints
Domain
Enterprise AI / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Kubernetes, distributed systems, Linux-based computing environments, multi-cloud infrastructure (GCP, Azure, AWS, OCI), compute/storage/network resource management, Golang, C++, GPU/TPU accelerator optimization
Preferred skills
None stated
Technologies
Kubernetes, GCP, Azure, AWS, OCI, Linux, Golang, C++, GPUs, TPUs
Responsibilities
Designing large, highly available distributed systems; deploying and operating AI platforms; troubleshooting complex technical challenges; managing compute/storage/network resources; creating customized deployments for customers
Seniority
Staff, hands-on IC