CareerPlanSign in

Principal AI Software Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-09-18 → 2026-09-25

Core

Lead full system software prototyping and workload characterization to evaluate hardware/software co-designed memory solutions for reducing TCO in AI inference workloads.

Role type

Principal AI Software Engineer (Systems & Memory Architecture)

Builds

Proof-of-concepts, evaluation frameworks, and performance models for memory-tiering architectures and KV Cache optimization.

Domain

AI Infrastructure / High-Performance Computing / Memory Systems

Deliverable

production ML models | infrastructure

Required skills

Linux kernel internals, memory management, I/O subsystems, NUMA, DMA, GPU/CPU/storage/network data paths, C/C++, Python, CUDA, distributed systems, workload characterization, system optimization

Preferred skills

Hardware/software co-design, CXL memory expansion, disaggregated prefill/decode architectures, multi-node cache-sharing topologies, inference runtime design

Technologies

NVIDIA GPU stacks (CUDA, NCCL, GDS, GDR), vLLM, SGLang, TensorRT-LLM, LMCache, SGLang HiCache, CXL, SSD-backed cache tiers

Responsibilities

Lead characterization and optimization of LLM inference workloads focusing on KV Cache capacity, placement, migration, and utilization; Develop software prototypes and instrumentation to evaluate memory-overcommit and offload techniques; Design and execute workload characterization studies for agentic, multi-turn, and long-context AI workloads; Analyze end-to-end data movement across GPU, CPU, storage, and networking subsystems; Build performance models to predict the impact of memory hierarchy innovations on large-scale deployments; Influence and shape hardware architecture and industry alignment over a three-to-six-year timeframe.

Seniority

Principal, strategy & mentorship

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.