Sr. Principal Software Engineer
Core
Optimize and deploy high-performance LLM inference pipelines for edge, embedded, and data center platforms to enable efficient AI companions in vehicles.
Role type
Senior Principal Software Engineer (LLM Inference Optimization)
Builds
Inference runtimes and optimized deployment pipelines for automotive AI products
Domain
Automotive AI / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
LLM inference optimization, CUDA kernel development, GPU architecture expertise, quantization strategies (INT8/INT4/FP4/FP8, AWQ, GPTQ), KV cache optimization, latency and throughput tuning
Preferred skills
Embedded systems deployment, custom CUDA kernel tuning, speculative decoding implementation
Technologies
vLLM, TensorRT-LLM, llama.cpp, QAIRT, CUDA
Responsibilities
Build and extend inference engines using custom CUDA kernels; implement quantization and memory layout optimizations; tune batching and decoding strategies for low latency; ensure efficient deployment on edge and embedded devices
Seniority
Senior, hands-on IC
