Member of Technical Staff (AI Inference Engineer)
Core
Build and run the inference engine behind every Perplexity query, deploying dozens of model architectures at scale with tight latency and cost budgets.
Role type
Senior IC AI Inference Engineer
Builds
High-throughput inference infrastructure for LLMs, retrieval, and multimodal models
Domain
AI/ML Systems, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models
Required skills
GPU programming (CUDA, Triton, CUTLASS), Rust systems programming, LLM architecture knowledge, distributed systems operation, performance profiling, API Gateway integration
Preferred skills
ML compilers (PyTorch internals, torch.compile), distributed GPU communication (NCCL, NVLink), low-precision inference (quantization), profiling tools (Nsight, CUDA-GDB), container orchestration (Kubernetes)
Technologies
Rust, Python, CUDA, CuTe DSL, PyTorch, JAX, TensorFlow, Kubernetes, NVLink, InfiniBand
Responsibilities
Support transformer-based models (retrieval, text-generation, multimodal) in inference infrastructure; Port CUDA kernels to CuTe DSL for GB200 and future hardware; Develop Rust-native serving runtime; Profile and fix bottlenecks in network ingress and GPU kernels; Build dashboards, alerts, and automated remediation for production incidents
Seniority
Senior, hands-on IC