AI Systems Administrator
Core
Implement and maintain a closed GPT environment with multiple LLM models, ensuring system health, resource allocation, and API integration for organizational use.
Role type
Senior AI Systems Administrator (Linux/GPU infrastructure)
Builds
Production AI/ML infrastructure, GPU-enabled LLM servers, and automated platform workflows
Domain
Defense & R&D, AI/ML Infrastructure, Linux Systems Administration
Deliverable
production ML models | infrastructure
Required skills
Linux administration (RHEL/Oracle), GPU driver management, Python scripting, Ansible automation, observability (Prometheus/Grafana), performance tuning, security operations
Preferred skills
Kubernetes, AWS/Azure, CI/Git workflows, root-cause analysis
Technologies
RHEL, Oracle Linux, Python, Ansible, Prometheus, Grafana, DCGM, NVML, Git, AWS, Azure
Responsibilities
Build and troubleshoot RHEL/Oracle systems supporting GPU workloads; Manage GPU enablement layer including driver/toolkit lifecycle; Implement observability for system, GPU, and storage performance; Couple observability with LLM performance to identify resource allocation issues; Maintain LLM servers for high uptime; Partner with storage/network peers to tune platform performance; Create automation for platform administration and Linux team workflows
Seniority
Senior, hands-on IC with leadership in platform redesign