CareerPlanSign in

Sr. Site Reliability Engineer

USA💼 Full-time💰 $165,000–$165,000🗓 2026-09-23 → 2026-09-26

Core

Manage GPU and CPU infrastructure deployments for Top Secret data centers and provide GPU-as-a-service support for external customers on bare-metal and virtualized platforms.

Role type

Senior Site Reliability Engineer (AI Infrastructure)

Builds

AI cluster solutions at 100,000+ GPU scale, on-premise Kubernetes and AI clusters, and distributed storage systems.

Domain

Defense/Security, AI Infrastructure, High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Linux system administration, Kubernetes cluster management, Infrastructure as Code (Terraform, Ansible), Containerization (OCI), Scripting (Bash, Python), Systems programming (Python, C++, Go), Database management, Monitoring and alerting, Team mentorship

Preferred skills

NVIDIA GPU deployment stacks, Distributed databases and data modeling, Large-scale server automation, TCP/IP networking, Cloud virtualization, Build and deployment systems (Bazel, Makefiles), Performance optimization

Responsibilities

Design, validate, and productize AI cluster solutions at 100,000+ GPU scale; Develop automation for deploying and managing on-premise Kubernetes and AI clusters; Deploy and manage databases, monitoring systems, and distributed storage; Support monitoring and alerting systems to maintain high availability; Identify reliability improvements and create innovative solutions for system availability; Mentor junior engineers and lead the team toward technical excellence.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.