Sr. Software Engineer, AI Infrastructure
Core
Design, deploy, and manage large-scale GPU and CPU infrastructure for AI clusters and data centers, providing GPU-as-a-service and ensuring system availability.
Role type
Senior Infrastructure Engineer (AI Systems)
Builds
Scalable AI clusters, GPU-as-a-service platforms, and on-premise Kubernetes environments
Domain
AI Infrastructure, High-Performance Computing, Data Center Operations
Deliverable
infrastructure
Required skills
Linux system administration, Kubernetes cluster management, Infrastructure as Code (Terraform/Ansible), Containerization (OCI), Scripting (Bash/Python), Distributed storage, Monitoring systems (via careerplan.io/jobs/22538-sr-software-engineer-ai-infrastructure-at-spacex)
Preferred skills
Python/C++/Go development, NVIDIA GPU stack deployment (Blackwell/Rubin), Performance optimization, TCP/IP networking, Cloud virtualization, Bazel/Makefiles
Responsibilities
Manage GPU and CPU infrastructure deployments to Top Secret data centers, Design and productize AI cluster solutions at 100,000-plus GPU scale, Develop automation for on-premise Kubernetes and AI clusters, Deploy and manage databases, monitoring systems, and distributed storage, Mentor and train junior engineers, Improve monitoring, alerting, and system availability
Seniority
Senior, hands-on IC with leadership responsibilities