AI Infrastructure Engineer
Core
Designing, implementing, and delivering production-grade software for high-performance, scalable inference systems powering Large Language Models (LLMs) and Vision-Language Models (VLMs) across cloud, edge, and hybrid environments.
Role type
Senior AI Inference Infrastructure Software Engineer
Builds
High-performance inference kernels, scalable serving systems, and optimized AIOS platform components for automotive applications.
Domain
Automotive AI, Large-scale Systems, Hardware Acceleration
Deliverable
production ML models
Required skills
C/C++, CUDA, PyTorch, Transformer architecture internals, kernel development, distributed inference systems, memory optimization, parallelism strategies, computer architecture, systems programming
Preferred skills
Hardware-aware model optimization, edge/embedded AI, inference serving system design, open source contributions
Technologies
CUDA, PyTorch, TensorFlow, GPU/NPU, NPU, DSP
Responsibilities
Design and implement scalable inference systems for LLMs/VLMs; Develop and optimize custom kernels for hardware accelerators; Integrate optimization techniques like KV-cache management and quantization; Partner with system/hardware teams for tight integration; Translate architectural requirements into production-ready software; Define evolution roadmap for LLM/VLM inference.
Seniority
Senior, hands-on IC