CareerPlanSign in

Sr. Principal Software Engineer

🌐 Remote💼 Full-time💰 $185,000–$185,000🗓 2026-06-30 → 2026-09-25

Core

Optimize and deploy high-performance LLM inference pipelines for edge, embedded, and data center platforms to enable efficient AI companions in vehicles.

Role type

Senior Principal Software Engineer (LLM Inference Optimization)

Builds

Inference runtimes and optimized deployment pipelines for automotive AI products

Domain

Automotive AI / Machine Learning Infrastructure

Deliverable

production ML models

Required skills

LLM inference optimization, CUDA kernel development, GPU architecture expertise, quantization strategies (INT8/INT4/FP4/FP8, AWQ, GPTQ), KV cache optimization, latency and throughput tuning

Preferred skills

Embedded systems deployment, custom CUDA kernel tuning, speculative decoding implementation

Technologies

vLLM, TensorRT-LLM, llama.cpp, QAIRT, CUDA

Responsibilities

Build and extend inference engines using custom CUDA kernels; implement quantization and memory layout optimizations; tune batching and decoding strategies for low latency; ensure efficient deployment on edge and embedded devices

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.