Site Reliability Engineer, AI Infrastructure (Starshield)
Core
Design, operate, and scale GPU/CPU infrastructure and AI clusters for critical national security missions and external customers.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
On-premise Kubernetes/AI clusters, GPU-as-a-Service platforms, and distributed storage systems for national security datacenters.
Domain
National Security / AI Infrastructure / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Infrastructure as Code (Terraform/Ansible), Containerization (Kubernetes/OCI), Scripting (Bash/Python), Systems programming (Python/C++/Go), Distributed systems, Networking (TCP/IP)
Preferred skills
Kubernetes cluster management, Linux boot process, CI/CD pipelines, Bazel/Makefiles, NVIDIA GPU deployment stacks (Blackwell/Rubin), Distributed databases
Technologies
Kubernetes, Terraform, Ansible, Python, C++, Go, Bash, Linux, NVIDIA GPUs, TCP/IP
Responsibilities
Manage GPU/CPU infrastructure deployments to Top Secret datacenters, Develop automation for on-premise Kubernetes/AI clusters, Monitor and alert on systems to ensure high availability, Collaborate with AI engineers to create scalable products, Identify and implement solutions for system availability improvements
Seniority
Senior, hands-on IC