CareerPlanSign in

Research Engineer - Model Evaluation & MLOps

San Francisco, CA💼 Full-time🗓 2026-08-31 → 2026-09-26

Core

Build tools and infrastructure to evaluate, deploy, and operate multimodal foundation models reliably, enabling rapid integration of internal and open-weight models on GPUs.

Role type

Research Engineer (Model Evaluation & MLOps)

Builds

Automated benchmarks for model quality and systems performance; CI/CD pipelines connecting research experiments to validated deployments; reusable tools for researchers to launch evaluations and compare experiments.

Domain

AI Infrastructure / Multimodal AI Models / GPU Systems

Deliverable

production ML models

Required skills

Python, PyTorch, TensorFlow, JAX, model evaluation, MLOps, experiment tracking, model versioning, GPU inference runtimes (vLLM, SGLang, TensorRT-LLM), containerized environments, debugging model workloads

Preferred skills

Hugging Face Transformers, AMD GPUs, ROCm, open-source contributions

Technologies

vLLM, SGLang, TensorRT-LLM, PyTorch, TensorFlow, JAX, Hugging Face

Responsibilities

Integrate new internal and open-weight language and multimodal models into GPU evaluation and inference environments; Build automated benchmarks for model quality and systems performance including latency, throughput, and memory usage; Build and maintain experiment tracking, model registry, and versioning for models, datasets, and evaluation configurations; Automate the path from research checkpoints to validated deployments through CI/CD and reproducible workflows; Profile end-to-end model workloads and collaborate with distributed systems, inference, and GPU kernel engineers on deeper performance issues.

Seniority

Mid-level, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.