CareerPlanSign in

Software Engineer, Data Infrastructure - Research

San Francisco💼 Full-time🗓 2025-09-18 → 2026-09-26

Core

Design and implement dataset infrastructure for OpenAI's next-generation LLM training stack, handling multimodal data and scaling pipelines across thousands of GPUs.

Role type

Senior IC data infrastructure engineer (LLM training)

Builds

Standardized dataset APIs, scale validation pipelines, and inspection tools for GPU-scale distributed training

Domain

AI research / Large Language Model training infrastructure

Deliverable

production ML models

Required skills

distributed systems, data pipelines, API design, modular code, scalable abstractions, debugging large-scale fleets, performance optimization

Preferred skills

data math, probability theory, distributed data theory, real-time dataset scaling

Technologies

GPU fleets, distributed training frameworks

Responsibilities

Design and maintain standardized dataset APIs for multimodal data; Build proactive testing and scale validation pipelines; Integrate datasets into training and inference pipelines; Document and maintain dataset interfaces; Establish safeguards for dataset reproducibility; Debug performance bottlenecks in distributed dataset loading; Provide visualization tools for dataset errors and bottlenecks

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.