Staff DevOps Engineer
Core
Design, build, and operate the infrastructure powering real-time AI inference across large-scale GPU fleets and global production systems.
Role type
Staff/Senior DevOps Engineer (Infrastructure & GPU Scale)
Builds
High-performance, automated infrastructure for AI inference serving
Domain
AI/ML Infrastructure, GPU Computing, Distributed Systems
Deliverable
infrastructure
Required skills
Linux system administration, Infrastructure-as-Code, CI/CD pipeline design, TCP/IP networking fundamentals, high-availability system operations, observability stack implementation, incident response and debugging, security hardening and compliance
Preferred skills
GPU infrastructure operations (NVIDIA drivers, CUDA), inference serving frameworks (vLLM, TensorRT, Triton), container runtime management, capacity planning for GPU workloads
Technologies
Linux, Kubernetes, Docker, vLLM, TensorRT, Triton, NVIDIA CUDA, TCP/IP, TLS, HTTP
Responsibilities
Provision and orchestrate GPU fleets and serverless/containerized production systems; automate infrastructure operations from provisioning to deployment safety; improve critical paths for request entrypoints, inference services, and networking; build observability backbones for capacity and issue detection; lead production operations, incident response, and post-incident improvements; strengthen security foundations through patching, secrets management, and access controls
Seniority
Staff/Senior, hands-on IC