Sr. Software Engineer, AI Infrastructure
Core
Design, validate, and productize AI cluster solutions at 100,000+ GPU scale; manage GPU and CPU infrastructure deployments in Top Secret data centers.
Role type
Senior IC infrastructure engineer (AI clusters)
Builds
Production AI clusters, GPU-as-a-service platforms, and automation for on-premise Kubernetes and AI environments (via careerplan.io/jobs/22537-sr-software-engineer-ai-infrastructure-at-spacex)
Domain
AI infrastructure, high-performance computing, secure data centers
Deliverable
production ML models
Required skills
Kubernetes, Linux, Terraform, Ansible, containerization (OCI), Bash, Python, C++, Go, distributed storage, monitoring systems
Preferred skills
Python development, Linux boot and systems configuration, continuous integration/build/deployment/monitoring, Bazel, Makefiles, performance optimization, distributed databases, TCP/IP networking, cloud virtualization, NVIDIA GPU deployment stacks
Responsibilities
Manage GPU and CPU infrastructure deployments; provide GPU-as-a-service support; design and validate AI cluster solutions; develop automation for deploying and managing on-premise Kubernetes and AI clusters; deploy and manage databases, monitoring systems, and distributed storage; mentor junior engineers
Seniority
Senior, hands-on IC