AI Inference Internship
Core
Intern working with the AI Inference team to optimize model serving latency and throughput for Perplexity products.
Role type
AI Inference Intern
Builds
Inference engine and deployments for single-node embeddings to distributed sparse Mixture-of-Experts models
Domain
AI / High-Performance Computing / Distributed Systems
Deliverable
infrastructure
Required skills
multi-threaded programming, networking, compilation, systems programming, ML frameworks (Torch, JAX), GPU programming (CUDA, Triton), High-Performance Computing (OpenMPI)
Responsibilities
Improve serving latency and throughput, bring up support for new models and inference optimizations, optimize inference across the entire stack from GPU kernels to serving endpoints
Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.