Research Engineer, Post-Training
Core
Build post-training pipelines to turn expert feedback and agent traces into models optimized for legal work.
Role type
Senior research engineer (post-training & agent optimization)
Builds
Production-grade post-training models and agent harnesses for legal domain
Domain
Legal technology + Large Language Models (LLMs) + Agentic AI
Deliverable
production ML models
Required skills
Post-training (SFT, RLHF/RLAIF, reward modeling, distillation), Python, experiment design, agent behavior analysis, failure mode identification
Preferred skills
Data/evaluation infrastructure, distributed training, GPU workloads, research publications
Technologies
Python, LLMs, agent frameworks
Responsibilities
Drive post-training experiments balancing cost, latency, and security; Optimize agent harnesses with domain-specific tools and retrieval strategies; Design grading and reward systems for high-stakes legal work; Analyze agent behavior to convert findings into training data or evals; Collaborate with researchers and partners on methodology and results
Seniority
Senior, hands-on IC