AI Research Engineer
Core
Optimizing deep learning models for deployment on edge AI platforms through model compression and efficient inference.
Role type
Senior IC AI Research Engineer (Model Optimization)
Builds
Optimized AI models for edge devices and NPUs
Domain
Edge AI, Hardware-Software Co-design, Model Compression
Deliverable
production ML models
Required skills
Deep learning, Model quantization (QAT, PTQ), Mixed precision optimization, Python, C++, CUDA, OpenCL, PyTorch, TensorFlow, ONNX Runtime, TVM, TensorRT, OpenVINO, Low-level hardware acceleration (SIMD, AVX, Tensor Cores, VNNI), Compiler optimizations (XLA, MLIR, LLVM)
Preferred skills
Knowledge distillation, Sparsity, Pruning, Model compression techniques, Benchmarking across hardware/software stacks
Technologies
PyTorch, TensorFlow, ONNX Runtime, TVM, TensorRT, OpenVINO, XLA, MLIR, LLVM, CUDA, OpenCL
Responsibilities
Research and develop quantization-aware training (QAT) and post-training quantization (PTQ) techniques; Implement low-bit precision optimizations (INT8, BF16); Design and optimize efficient inference algorithms for latency, memory footprint, and power efficiency; Collaborate with hardware engineers to optimize model execution for edge devices and NPUs; Analyze accuracy trade-offs and develop calibration techniques for quantized models; Benchmark performance across different hardware and software stacks.
Seniority
Senior, hands-on IC