CareerPlanSign in

Site Reliability Engineer, AI Infrastructure (Starshield)

Hawthorne, CA💼 Full-time💰 $1–$1🗓 2026-09-15 → 2026-09-26

Core

Design, operate, and scale GPU/CPU infrastructure and AI clusters for critical national security missions and external customers.

Role type

Senior Site Reliability Engineer (AI Infrastructure)

Builds

On-premise Kubernetes/AI clusters, GPU-as-a-Service platforms, and distributed storage systems for national security datacenters.

Domain

National Security / AI Infrastructure / High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Linux system administration, Infrastructure as Code (Terraform/Ansible), Containerization (Kubernetes/OCI), Scripting (Bash/Python), Systems programming (Python/C++/Go), Distributed systems, Networking (TCP/IP)

Preferred skills

Kubernetes cluster management, Linux boot process, CI/CD pipelines, Bazel/Makefiles, NVIDIA GPU deployment stacks (Blackwell/Rubin), Distributed databases

Technologies

Kubernetes, Terraform, Ansible, Python, C++, Go, Bash, Linux, NVIDIA GPUs, TCP/IP

Responsibilities

Manage GPU/CPU infrastructure deployments to Top Secret datacenters, Develop automation for on-premise Kubernetes/AI clusters, Monitor and alert on systems to ensure high availability, Collaborate with AI engineers to create scalable products, Identify and implement solutions for system availability improvements

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.