CareerPlanSign in

Inference Intern

San Jose💼 Internship🗓 2025-12-08 → 2026-09-26

Core

Designing next-generation AI accelerators and optimizing compute architectures for inference workloads.

Role type

Inference Architecture Intern

Builds

AI accelerator chips, racks, and software for inference workloads

Domain

Hardware/Software co-design for AI inference

Deliverable

production ML models

Required skills

Python, C++, Linux internals, GPU/TPU architecture, compiler knowledge, high-speed interconnects (NVLink, InfiniBand), transformer model architectures, inference serving stacks (vLLM, SGLang)

Preferred skills

Rust, kernel-level/user-space networking, distributed systems concepts, Mixture-of-Experts (MoE), SIMD optimizations, PyTorch, JAX

Technologies

vLLM, SGLang, PyTorch, JAX, NVLink, InfiniBand

Responsibilities

Port state-of-the-art models to architecture, build runtime for multi-node inference, optimize routing and communication layers, profile and debug performance bottlenecks, co-design HW instructions and model operations, implement high-performance software components for Model Toolkit

Seniority

Intern

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.