LLM Engineer (Data Generation)
Core
Design, generate, evaluate, and improve training data to enhance Large Language Model (LLM) performance from a data-centric AI perspective.
Role type
Senior IC LLM Engineer (Data Generation)
Builds
Instruction, Preference, Reasoning, and Domain-specific datasets; automated data generation pipelines.
Domain
Generative AI, Large Language Models, Natural Language Processing
Deliverable
production ML models
Required skills
LLM and Machine Learning expertise, Python development, data pipeline automation, synthetic data generation, data evaluation strategies, prompt engineering, data curation
Preferred skills
LLM pre-training and fine-tuning (SFT, DPO, RLHF), LLM evaluation framework development, complex data design (multi-turn, agent, reasoning), large-scale data processing environments (Airflow, Ray, Spark), open-source contributions
Technologies
Python, LLM APIs, Prompting frameworks, Data Evaluation tools (OpenAI Evals, DeepEval), Workflow orchestration tools (Airflow, Ray, Spark)
Responsibilities
Analyze model performance bottlenecks and failure cases to define data requirements; design and generate diverse training data; build and operate automated data generation pipelines; define and execute data quality and evaluation metrics; iteratively improve data generation strategies based on experimental results.
Seniority
Senior, hands-on IC