AI Infrastructure & Accelerator Architect
Core
Architecting high-performance AI kernels, custom operations, and hardware-software co-design for TPU and GPU architectures to optimize training and inference efficiency.
Role type
Principal AI Infrastructure & Accelerator Architect
Builds
Enterprise-grade benchmarking suites, automated autotuning frameworks, regression analysis pipelines, and foundational infrastructure for AI training and serving.
Domain
AI Infrastructure, High-Performance Computing, Compiler Engineering
Deliverable
production ML models | infrastructure
Required skills
Distributed systems architecture, C++/Python system design, Kernel-level performance optimization, Hardware accelerator architecture (TPU/GPU), Compiler fundamentals (MLIR, LLVM, OpenXLA), Model quantization, Low-precision arithmetic, Heterogeneous compute, Multi-node scale-out fabrics
Preferred skills
JAX/PyTorch framework expertise, Attention mechanisms, Mixture of Experts, Code generation toolchains, Open-source library scaling, Strategic communication
Technologies
CUDA, TPU, GPU, LLVM, MLIR, OpenXLA, JAX, PyTorch, Python, NodeJS
Responsibilities
Define multi-year technical roadmap for high-performance AI kernels and hardware-software co-design; Scale and mentor a world-class technical practice by setting architectural governance; Act as principal technical liaison with ML researchers and compiler teams to remove bottlenecks; Architect foundational infrastructure including benchmarking suites and autotuning frameworks; Track industry shifts in hardware and compiler innovations to improve efficiency.
