Staff Deep Learning Compiler Engineer
Core
Design and optimize the compiler stack (MLIR, TVM, LLVM) to bridge neural network frameworks to Quadric's proprietary General Purpose Neural Processing Unit (GPNPU) architecture for edge AI workloads.
Role type
Staff Deep Learning Compiler Engineer
Builds
Optimized low-level kernel code and compiler infrastructure for edge AI devices
Domain
Edge AI, Hardware-Software Co-Design, Compiler Infrastructure
Deliverable
production ML models
Required skills
Deep learning compiler development, C++ (14/17/20), Python, MLIR, Apache TVM, LLVM, XLA, TensorRT, graph-level optimization, operator fusion, model quantization, memory hierarchy management, parallel processing
Preferred skills
Low-level kernel optimization, SIMD/vector programming, FPGA bring-up, hardware emulation, cycle-accurate simulator development, post-training/QAT optimization
Technologies
MLIR, Apache TVM, LLVM, PyTorch, TensorFlow, ONNX, C++, Python
Responsibilities
Design and implement compiler optimization passes targeting proprietary processor architecture; Develop lowering pathways from high-level ML frameworks to optimized low-level kernel code; Implement graph-level optimizations including operator fusion and layout transformation; Collaborate with hardware teams to define instruction set extensions and acceleration features; Develop software simulators and test benches to validate compiler correctness; Benchmark and profile end-to-end model performance to resolve compiler bottlenecks
Seniority
Staff, hands-on IC with strategic influence