Software Engineer, Data Infrastructure - Research
Core
Design and implement dataset infrastructure for OpenAI's next-generation LLM training stack, handling multimodal data and scaling pipelines across thousands of GPUs.
Role type
Senior IC data infrastructure engineer (LLM training)
Builds
Standardized dataset APIs, scale validation pipelines, and inspection tools for GPU-scale distributed training
Domain
AI research / Large Language Model training infrastructure
Deliverable
production ML models
Required skills
distributed systems, data pipelines, API design, modular code, scalable abstractions, debugging large-scale fleets, performance optimization
Preferred skills
data math, probability theory, distributed data theory, real-time dataset scaling
Technologies
GPU fleets, distributed training frameworks
Responsibilities
Design and maintain standardized dataset APIs for multimodal data; Build proactive testing and scale validation pipelines; Integrate datasets into training and inference pipelines; Document and maintain dataset interfaces; Establish safeguards for dataset reproducibility; Debug performance bottlenecks in distributed dataset loading; Provide visualization tools for dataset errors and bottlenecks
Seniority
Senior, hands-on IC
