Member of Technical Staff - Data Flywheel Infra, Frontier Models
Core
Build scalable infrastructure for ingesting, processing, curating, and serving first-party and third-party data to support pre-training and post-training of frontier models, while establishing governance and policy enforcement within the data platform.
Role type
Senior IC data infrastructure engineer (AI training data flywheel)
Builds
Data flywheel infrastructure, automated data acquisition pipelines, evaluation-to-data feedback loops, and synthetic data generation systems.
Domain
AI/ML infrastructure, data governance, privacy-preserving computing, and large language model training workflows.
Deliverable
production ML models
Required skills
Data engineering, data modeling, AI training-data governance, privacy-aware system design, data lifecycle management, synthetic data generation, evaluation feedback loop design, multimodal data handling, LLM training workflows, data quality metrics, policy enforcement.
Preferred skills
Experience with data clean rooms, confidential computing, model graders, reward signals, active learning, data-mixture optimization, agentic datasets.
Technologies
Data pipelines, versioning systems, PII detection tools, access control frameworks, synthetic data generators, evaluation frameworks.
Responsibilities
Design and build scalable systems for data ingestion, processing, and serving for model training; implement governance controls for data provenance, licensing, consent, and PII protection; develop automated pipelines for data classification, filtering, and enrichment; create feedback loops connecting model evaluations to data acquisition; build metrics for dataset quality and policy compliance.
Seniority
Senior, hands-on IC