ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)
Core
Building and deploying reinforcement learning (RL) training environments, synthetic data pipelines, and fine-tuning jobs for open-source software engineering LLMs (SWE models) used in self-serve and enterprise products.
Role type
Senior ML Systems Engineer (LLM Training & Infrastructure)
Builds
Production-grade SWE models, RL training infrastructure, and synthetic data pipelines for code generation agents.
Domain
Artificial Intelligence / Software Engineering / Large Language Models
Deliverable
production ML models
Required skills
Python, Go, PyTorch, Docker, Kubernetes, Cloud platforms (GCP/AWS/Azure), Data engineering, Custom training loops, RL objectives, Evaluation frameworks
Preferred skills
Synthetic data generation, SQL, Apache Iceberg, DuckDB, Distributed LLM training, Reward shaping, LLM-as-a-judge, Open-source contributions
Technologies
PyTorch, Docker, Kubernetes, GCP, AWS, Azure, SQL, Apache Iceberg, DuckDB
Responsibilities
Develop and manage synthetic data generation pipelines for RL fine-tunes; Design, build, and deploy containerized services for RL infrastructure; Build and iterate on large-scale RL loops for code writing and testing; Architect synthetic data pipelines and deploy using containerization; Improve evaluation suites for code models and analyze failure modes.
Seniority
Senior, hands-on IC