Staff Software Engineer , Anywhere Cloud - AI Systems & Runtimes
Core
Building the architecture and delivery of a cloud-native AI platform that bridges cutting-edge AI research with production-grade Kubernetes environments to enable agentic AI and seamless model consumption.
Role type
Staff Software Engineer (AI Systems & Runtimes)
Builds
Production-grade AI services, inference servers, and orchestration patterns for LLMs on Kubernetes.
Domain
Cloud infrastructure, Kubernetes, Large Language Models (LLMs), AI Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Python, Go or Rust/C++, LLM deployment and runtimes (vLLM, Triton), Kubernetes (CRDs, Operators), RAG pipelines, GPU resource scheduling
Preferred skills
Model fine-tuning (PEFT, LoRA), CUDA programming, Open source contributions
Technologies
Kubernetes, KServe, KubeRay, Knative, vLLM, Triton, LangChain, LlamaIndex, Docker, Go, Node.js, Python, CUDA
Responsibilities
Design and implement scalable application services wrapping AI capabilities; Lead deployment of inference servers using K8s-native patterns; Build internal tooling and SDKs for AI integration; Architect RAG pipelines and prompt management services; Ensure AI workloads are secure and optimized for GPU scheduling; Partner with UI/UX and Product teams on platform usability.
Seniority
Staff, hands-on IC with architectural leadership