CareerPlanGet AI match score →

AI Engineer - Model Adaptation & Evaluation

💼 Full-time🗓 2026-06-25

Core

Adapt, evaluate, and deploy open-weight language models for specialized operational use cases in constrained or secure environments.

Role type

LLM Fine-Tuning Engineer

Builds

Task-specific systems derived from foundation models

Domain

Artificial Intelligence / Large Language Models

Deliverable

production ML models

Required skills

supervised fine-tuning (SFT), parameter-efficient fine-tuning (LoRA, QLoRA), preference tuning (DPO), dataset curation, synthetic data generation, PyTorch, Hugging Face Transformers, GPU optimization, quantization-aware workflows

Preferred skills

debugging model regressions, translating domain needs into evaluation criteria

Technologies

PyTorch, Hugging Face Transformers

Responsibilities

Fine-tune and adapt open-weight LLMs for specialized use cases; Design, run, and compare post-training approaches; Build, clean, and improve high-quality datasets; Generate and validate synthetic data; Define and implement evaluation methods; Analyze model regressions and failure modes; Package, serve, and operationalize tuned models in restricted environments

Seniority

Mid-to-Senior, hands-on IC

Rewrite
## About the role We are looking for an LLM Fine-Tuning Engineer to help adapt, evaluate, and deploy open-weight language models for specialized operational use cases. The role focuses on the full post-training lifecycle: dataset design, supervised fine-tuning, parameter-efficient fine-tuning, synthetic data generation, model evaluation, and deployment support. You will work closely with engineering and domain teams to transform foundation models into reliable, task-specific systems that can operate effectively in constrained, local, or secure environments. This is not a generic prompt engineering role. We are looking for someone with hands-on experience in model adaptation, training data quality, evaluation methodology, and practical deployment constraints. ## What You'll Do - Fine-tune and adapt open-weight LLMs for specialized internal or customer-facing use cases. - Design, run, and compare different post-training approaches, including supervised fine-tuning (SFT), parameter-efficient fine-tuning such as LoRA or QLoRA, preference tuning approaches such as DPO where appropriate, and full fine-tuning when justified by the task and infrastructure constraints. - Build, clean, and improve high-quality datasets for training, validation, and evaluation. - Generate and validate synthetic data to expand coverage, improve robustness, and accelerate iteration cycles. - Define and implement evaluation methods that go beyond generic benchmark scores. - Analyze model regressions, failure modes, and behavioral changes after fine-tuning. - Work with engineering teams to package, serve, and operationalize tuned models in local, private, or restricted environments. - Collaborate with domain experts to translate real-world needs into measurable model requirements and evaluation criteria. ## What We're Looking For - Strong practical experience fine-tuning LLMs or adjacent foundation models in production, applied research, or serious experimental environments. - Hands-on experience with multiple post-training techniques, ideally including SFT, LoRA, and QLoRA. - Experience building, curating, and maintaining datasets for model training and evaluation. - Experience generating, filtering, and validating synthetic data for model improvement. - Strong Python skills and practical experience with PyTorch, Hugging Face Transformers, tokenization, preprocessing, and training pipelines. - Good understanding of GPU constraints, memory/performance trade-offs, quantization-aware workflows, and training optimization. - Strong debugging mindset, with the ability to understand why a model improved, regressed, overfit, or failed.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗