CareerPlanGet AI match score →

AI Inference Internship

London, England, UK💼 Internship🗓 2026-07-16 → 2026-07-19

Core

Intern working with the AI Inference team to optimize model serving latency and throughput for Perplexity products.

Role type

AI Inference Intern

Builds

Inference engine and deployments for single-node embeddings to distributed sparse Mixture-of-Experts models

Domain

AI / High-Performance Computing / Distributed Systems

Deliverable

infrastructure

Required skills

multi-threaded programming, networking, compilation, systems programming, ML frameworks (Torch, JAX), GPU programming (CUDA, Triton), High-Performance Computing (OpenMPI)

Responsibilities

Improve serving latency and throughput, bring up support for new models and inference optimizations, optimize inference across the entire stack from GPU kernels to serving endpoints

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗