Senior Deep Learning Compiler Engineer - HW-SW Codesign
Core
Develop compiler optimization algorithms for deep learning workloads, optimizing inference and training performance for the JAX framework and OpenXLA compiler on NVIDIA GPUs.
Role type
Senior IC deep learning compiler engineer (HW-SW codesign)
Builds
High-performance, production-grade AI software for next-generation AI systems
Domain
AI computing, deep learning frameworks, GPU architecture
Deliverable
production ML models
Required skills
compiler optimization, graph partitioning, tensor sharding, performance tuning, code generation, C/C++ programming, high-performance computing, distributed programming, GPU architecture knowledge
Preferred skills
CUDA/OpenCL programming, deep learning framework design, open-source compiler experience (XLA, TVM, MLIR, LLVM, Triton), mentoring
Technologies
JAX, OpenXLA, MLIR, LLVM, OpenAI Triton, CUDA, OpenCL, GPU hardware
Responsibilities
Craft and implement compiler optimization techniques for deep learning network graphs; Design novel graph partitioning and tensor sharding techniques; Perform performance tuning and analysis; Generate code for NVIDIA GPU backends; Design user-facing features in JAX and related libraries; Collaborate with hardware engineering teams to design AI compiler software features
Seniority
Senior, hands-on IC