AI Frameworks Engineer – GPU Performance for Generative AI (OpenVINO)
Core
Building high-performance, HW-aware software for efficient execution of generative AI models (LLMs, diffusion) on Intel GPU architectures.
Role type
Senior IC systems software engineer (GPU performance)
Builds
Production inference runtimes and optimized kernels for Intel GPUs
Domain
AI inference / GPU hardware / Systems programming
Deliverable
production ML models
Required skills
C++, C, Python, GPU architecture, parallel computing, system-level software design, performance optimization, debugging complex distributed systems
Preferred skills
SIMD programming, accelerator programming models, generative AI system perspective, AI runtime frameworks, data structures and algorithms
Technologies
OpenVINO, Intel GPUs, C++, CUDA (implied by GPU context), Python
Responsibilities
Take ownership of performance-critical paths for generative AI workloads; Analyze end-to-end execution to identify compute, memory, and bandwidth bottlenecks; Implement and optimize generative AI techniques for Intel GPU architectures; Translate deep understanding of GPU hardware into efficient software designs; Diagnose and resolve issues spanning runtime, kernel, driver, and hardware boundaries
Seniority
Senior, hands-on IC