CareerPlanSign in

Inference Software Engineer

San Jose💼 Full-time🗓 2025-06-17 → 2026-09-25

Core

Building and optimizing runtime systems for multi-node inference on custom hardware architectures to support state-of-the-art AI models.

Role type

Senior Inference Software Engineer (System/Runtime)

Builds

Multi-node inference runtime, programming abstractions, and testing capabilities for frontier intelligence hardware.

Domain

AI Infrastructure / High-Performance Computing / Custom Accelerator Hardware

Deliverable

production ML models

Required skills

C++, Rust, Linux internals, GPU/TPU architecture, high-speed interconnects (NVLink, InfiniBand), PyTorch, JAX, performance profiling, distributed systems

Preferred skills

Kernel-level networking stacks, consensus protocols, Transformer/MoE architectures, SIMD optimizations

Technologies

C++, Rust, PyTorch, JAX, NVLink, InfiniBand, Linux

Responsibilities

Port state-of-the-art models to custom architecture, build and scale runtime for multi-node inference, optimize routing and communication layers, debug performance bottlenecks

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.