UK Internship Program
Core
Intern working with the AI Inference team to optimize model serving latency and throughput for Perplexity's conversational answer engine.
Role type
AI Inference Intern
Builds
Inference engine and deployments for conversational AI models
Domain
AI / Machine Learning / High-Performance Computing
Deliverable
production ML models
Required skills
multi-threaded programming, networking, systems programming, ML frameworks (Torch, JAX), GPU programming (CUDA, Triton), High-Performance Computing (OpenMPI)
Preferred skills
performance-related subjects (HPC, Compilers, Distributed Systems)
Technologies
Torch, JAX, CUDA, Triton, OpenMPI
Responsibilities
Improve serving latency and throughput, bring up support for new models and inference optimizations, optimize inference across the entire stack from GPU kernels to serving endpoints
Seniority
Intern
Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.