Principal Software Engineering Manager - Substrate efficiency
Core
Lead engineering team to optimize inference runtime efficiency, model execution performance, and throughput per GPU for large-scale AI/ML systems.
Role type
Principal Software Engineering Manager (Inference Runtime)
Builds
High-performance inference engines and distributed systems for AI/ML workloads
Domain
Artificial Intelligence / Machine Learning / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
C, C++, C#, Java, JavaScript, Python, distributed systems, system throughput optimization, resource utilization, systems thinking, workload scheduling, batching, infrastructure efficiency, AI/ML inference systems, GPU-based workloads, runtime optimization, cost-per-query optimization
Preferred skills
N/A
Technologies
C, C++, C#, Java, JavaScript, Python
Responsibilities
Define and drive strategy to improve throughput per GPU through runtime optimizations; Establish metrics, telemetry, and experimentation frameworks to measure efficiency gains; Own live-site performance, reliability, and operational excellence for inference engines at scale; Drive alignment across partner teams on engine interfaces, performance goals, and optimization priorities; Increase engineering agility for faster experimentation and rollout of performance improvements.
Seniority
Principal, hands-on IC with management