Principal Research Scientist, Synthetic Data Generation
Core
Define technical direction and build open-source libraries for synthetic data generation to feed pre- and post-training of large language models (LLMs) like Nemotron.
Role type
Principal Research Scientist (Synthetic Data Generation)
Builds
Open-source libraries, SDKs, and scalable data generation pipelines for text, code, structured, and multimodal data.
Domain
Generative AI, Large Language Models, Multimodal Machine Learning
Deliverable
production ML models
Required skills
Generative modeling, LLM training pipelines, automated quality evaluation, differential privacy, open-source library development, distributed inference optimization, research publication
Preferred skills
Agentic AI and tool-use training, reinforcement learning environment design, multimodal generation (vision-language, video, audio), regulated industry data synthesis
Technologies
LLMs, vLLM, TGI, Git, CI/CD
Responsibilities
Build and scale data generation pipelines using LLM-based methods; pioneer data generation for agentic and tool-use training; advance multimodal synthetic data generation; develop privacy-preserving synthesis techniques; maintain open-source libraries with clean APIs; drive software excellence with modern tooling; publish original research at top conferences; mentor scientists and engineers.
Seniority
Principal, strategy & mentorship