Inference Software Engineer
Core
Building and optimizing runtime systems for multi-node inference on custom hardware architectures to support state-of-the-art AI models.
Role type
Senior Inference Software Engineer (System/Runtime)
Builds
Multi-node inference runtime, programming abstractions, and testing capabilities for frontier intelligence hardware.
Domain
AI Infrastructure / High-Performance Computing / Custom Accelerator Hardware
Deliverable
production ML models
Required skills
C++, Rust, Linux internals, GPU/TPU architecture, high-speed interconnects (NVLink, InfiniBand), PyTorch, JAX, performance profiling, distributed systems
Preferred skills
Kernel-level networking stacks, consensus protocols, Transformer/MoE architectures, SIMD optimizations
Technologies
C++, Rust, PyTorch, JAX, NVLink, InfiniBand, Linux
Responsibilities
Port state-of-the-art models to custom architecture, build and scale runtime for multi-node inference, optimize routing and communication layers, debug performance bottlenecks
Seniority
Senior, hands-on IC
