Member of Technical Staff - Compute Infrastructure
Core
Building one of the world's largest AI supercomputers from the ground up, owning both the raw GPU supercomputer and the platform layer to accelerate Grok's training speed and AI progress.
Role type
Senior IC compute infrastructure engineer (GPU supercomputing)
Builds
Massive GPU clusters, low-level GPU kernels, Linux kernel internals, custom container orchestration, and distributed systems for extreme-scale training and inference.
Domain
AI infrastructure / High-performance computing / GPU systems
Deliverable
production ML models
Required skills
Low-level systems programming (C/C++ or Rust), GPU kernel optimization (CUDA, CUTLASS, Tensor Cores), Linux kernel internals, distributed systems design, profiling and debugging at cluster scale, infrastructure-as-code automation
Preferred skills
Experience building exabyte-scale storage systems, large-scale GPU cluster operations, virtualization (KVM, Firecracker), reasoning from first principles for memory-bound and compute-bound scenarios
Technologies
CUDA, CUTLASS, Nsight, KVM, Firecracker, Kubernetes, Linux
Responsibilities
Design and optimize massive GPU clusters for extreme-scale workloads; Develop and tune low-level CUDA kernels; Work on Linux kernel internals and resource isolation; Build custom container orchestration and virtualization layers; Profile and eliminate bottlenecks across GPU memory hierarchy and networking; Create and maintain infrastructure-as-code and automation tools; Collaborate with AI research teams to deliver production-grade performance
Seniority
Senior, hands-on IC