Sr. Site Reliability Engineer
Core
Manage GPU and CPU infrastructure deployments for Top Secret data centers and provide GPU-as-a-service support for external customers on bare-metal and virtualized platforms.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
AI cluster solutions at 100,000+ GPU scale, on-premise Kubernetes and AI clusters, and distributed storage systems.
Domain
Defense/Security, AI Infrastructure, High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Kubernetes cluster management, Infrastructure as Code (Terraform, Ansible), Containerization (OCI), Scripting (Bash, Python), Systems programming (Python, C++, Go), Database management, Monitoring and alerting, Team mentorship
Preferred skills
NVIDIA GPU deployment stacks, Distributed databases and data modeling, Large-scale server automation, TCP/IP networking, Cloud virtualization, Build and deployment systems (Bazel, Makefiles), Performance optimization
Responsibilities
Design, validate, and productize AI cluster solutions at 100,000+ GPU scale; Develop automation for deploying and managing on-premise Kubernetes and AI clusters; Deploy and manage databases, monitoring systems, and distributed storage; Support monitoring and alerting systems to maintain high availability; Identify reliability improvements and create innovative solutions for system availability; Mentor junior engineers and lead the team toward technical excellence.
Seniority
Senior, hands-on IC with mentorship responsibilities