AI Runtime Engineer
Core
Develop and optimize the execution stack for next-generation AI accelerators to enable efficient deep learning model inference and training on specialized hardware.
Role type
AI Runtime Engineer
Builds
Low-latency, high-performance runtime software for AI inference and training across cloud and edge environments.
Domain
AI hardware systems, edge-to-cloud computing, deep learning execution
Deliverable
production ML models
Required skills
C/C++, low-level systems programming, task scheduling, memory management, kernel execution strategies, data movement optimization (PCIe, DMA), graph execution optimizations, AI compiler integration, debugging and profiling
Preferred skills
Hardware-aware optimizations, dataflow architectures, deep learning execution frameworks, AI model deployment pipelines
Technologies
OpenVino, ONNX Runtime, vLLM, LLVM, MLIR, XLA, TVM, TensorRT, Triton, TensorFlow Serving
Responsibilities
Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators; Implement task scheduling, memory management, and kernel execution strategies; Optimize data movement between host and device; Design and implement high-performance APIs for AI Inference frameworks; Work on graph execution optimizations; Integrate runtime components with AI compilers; Ensure scalability and reliability of the AI runtime
Seniority
Mid-level, hands-on IC