Senior Scientist, Synthetic Data Generation
Core
Building synthetic data generation pipelines and open-source libraries to create datasets for pre- and post-training of large language models (LLMs) and multimodal AI systems.
Role type
Senior Scientist (Applied Research & Software Engineering)
Builds
Synthetic datasets (text, code, structured, multimodal) and open-source SDKs/libraries for NVIDIA NeMo ecosystem.
Domain
Generative AI, Large Language Models, Multimodal Machine Learning
Deliverable
production ML models
Required skills
LLM architecture and training, generative modeling, multimodal generation (vision-language, audio, video), software library development, automated quality evaluation, distributed inference optimization, research publication
Preferred skills
Open-source contributions, agentic RL post-training, scalable data pipeline architecture
Technologies
vLLM, TGI, NVIDIA NeMo, Git, CI/CD
Responsibilities
Build synthetic data pipelines using LLM-based methods, advance multimodal generation capabilities, design and maintain open-source libraries with clean APIs, drive software excellence with modern tooling, publish original research at top conferences, mentor interns and junior researchers
Seniority
Senior, hands-on IC with research leadership