Inference Intern
Core
Designing next-generation AI accelerators and optimizing compute architectures for inference workloads.
Role type
Inference Architecture Intern
Builds
AI accelerator chips, racks, and software for inference workloads
Domain
Hardware/Software co-design for AI inference
Deliverable
production ML models
Required skills
Python, C++, Linux internals, GPU/TPU architecture, compiler knowledge, high-speed interconnects (NVLink, InfiniBand), transformer model architectures, inference serving stacks (vLLM, SGLang)
Preferred skills
Rust, kernel-level/user-space networking, distributed systems concepts, Mixture-of-Experts (MoE), SIMD optimizations, PyTorch, JAX
Technologies
vLLM, SGLang, PyTorch, JAX, NVLink, InfiniBand
Responsibilities
Port state-of-the-art models to architecture, build runtime for multi-node inference, optimize routing and communication layers, profile and debug performance bottlenecks, co-design HW instructions and model operations, implement high-performance software components for Model Toolkit
Seniority
Intern
