Principal Software Engineer, AI Inference Runtime
Core
Define architecture and roadmap for distributed AI inference runtime, optimizing scheduling, batching, memory management, and kernel performance for SOTA models.
Role type
Principal Software Engineer (AI Inference Runtime)
Builds
Distributed inference runtime components, optimized kernels, and production validation systems for Arm's AI platform.
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
C++, Rust, Python, concurrency, parallel programming, memory management, data movement, profiling, debugging, kernel optimization (via careerplan.io/jobs/100399715424-principal-software-engineer-ai-inference-runtime-at-arm)
Preferred skills
Inference schedulers, cache managers, batching systems, accelerator programming, assembly, intrinsics, model parallelism, collective communication, graph optimization
Technologies
C++, Rust, Python, Arm hardware, distributed systems
Responsibilities
Define architecture and roadmap for AI inference runtime; Enable new model architectures via operator support and optimization; Profile system bottlenecks and develop optimized kernels; Evaluate inference techniques and build benchmarking/validation systems; Partner with cross-functional teams and mentor engineers.
Seniority
Principal, hands-on IC with strategy & mentorship
