CareerPlanGet AI match score →

Inference Technical Lead, Sora

San Francisco💼 Full-time🗓 2025-04-21 → 2026-07-31

Core

Optimizing GPU inference performance and scalability for OpenAI's Sora multimodal foundation model.

Role type

Senior IC GPU Inference Engineer

Builds

High-throughput, reliable model serving infrastructure for multimodal AI products

Domain

Generative AI / Multimodal Systems

Deliverable

production ML models

Required skills

GPU inference optimization, kernel-level systems programming, data movement optimization, low-level performance tuning, distributed systems design

Preferred skills

Model design for inference efficiency, scaling high-performance AI systems, navigating technical ambiguity

Technologies

CUDA, PyTorch, Kubernetes, gRPC

Responsibilities

Optimize model serving efficiency and system throughput, drive kernel and data movement optimizations, partner with research teams on inference-friendly model design, build and improve critical serving infrastructure

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗