CareerPlanSign in

Principal Software Engineer, AI Inference Runtime

Seattle, Washington💼 Full-time💰 $262,700–$262,700🗓 2026-09-08 → 2026-10-07
Hybrid

Core

Define architecture and roadmap for distributed AI inference runtime, optimizing scheduling, batching, memory management, and kernel performance for SOTA models.

Role type

Principal Software Engineer (AI Inference Runtime)

Builds

Distributed inference runtime components, optimized kernels, and production validation systems for Arm's AI platform.

Domain

AI Infrastructure / High-Performance Computing

Deliverable

production ML models

Required skills

C++, Rust, Python, concurrency, parallel programming, memory management, data movement, profiling, debugging, kernel optimization (via careerplan.io/jobs/100399715424-principal-software-engineer-ai-inference-runtime-at-arm)

Preferred skills

Inference schedulers, cache managers, batching systems, accelerator programming, assembly, intrinsics, model parallelism, collective communication, graph optimization

Technologies

C++, Rust, Python, Arm hardware, distributed systems

Responsibilities

Define architecture and roadmap for AI inference runtime; Enable new model architectures via operator support and optimization; Profile system bottlenecks and develop optimized kernels; Evaluate inference techniques and build benchmarking/validation systems; Partner with cross-functional teams and mentor engineers.

Seniority

Principal, hands-on IC with strategy & mentorship