Staff System Engineer – AI Infrastructure
Core
Design and validate hands-on engineering solutions for AI workloads across heterogeneous GPU environments, on-premises datacenters, private, and public-cloud infrastructures.
Role type
Staff System Engineer (AI Infrastructure)
Builds
Validated technical solutions, scalable product capabilities, and infrastructure tooling for AI inference and serving.
Domain
Data & AI Platform, Cloud Infrastructure, GPU Computing
Deliverable
production ML models | infrastructure
Required skills
Systems software engineering, distributed infrastructure, performance engineering, Linux, containers, GPU runtime optimization, root cause analysis, prototyping, benchmarking, heterogeneous environment management
Preferred skills
AI inference and serving technologies (NVIDIA NIM, vLLM, SGLang, Triton), Kubernetes, GPU resource management (MIG), high-performance GPU networking diagnostics
Technologies
Linux, Kubernetes, NVIDIA GPUs, containers, public cloud, private cloud, on-premises datacenters
Responsibilities
Investigate and design solutions for AI workloads across diverse environments; develop and optimize infrastructure for production AI workloads including inference and serving; diagnose and resolve performance, reliability, and scalability issues across the full stack; build diagnostic and validation tooling; translate infrastructure findings into product improvements.
Seniority
Staff, hands-on IC with strategic impact