Research Engineer - Model Evaluation & MLOps
Core
Build tools and infrastructure to evaluate, deploy, and operate multimodal foundation models reliably, enabling rapid integration of internal and open-weight models on GPUs.
Role type
Research Engineer (Model Evaluation & MLOps)
Builds
Automated benchmarks for model quality and systems performance; CI/CD pipelines connecting research experiments to validated deployments; reusable tools for researchers to launch evaluations and compare experiments.
Domain
AI Infrastructure / Multimodal AI Models / GPU Systems
Deliverable
production ML models
Required skills
Python, PyTorch, TensorFlow, JAX, model evaluation, MLOps, experiment tracking, model versioning, GPU inference runtimes (vLLM, SGLang, TensorRT-LLM), containerized environments, debugging model workloads
Preferred skills
Hugging Face Transformers, AMD GPUs, ROCm, open-source contributions
Technologies
vLLM, SGLang, TensorRT-LLM, PyTorch, TensorFlow, JAX, Hugging Face
Responsibilities
Integrate new internal and open-weight language and multimodal models into GPU evaluation and inference environments; Build automated benchmarks for model quality and systems performance including latency, throughput, and memory usage; Build and maintain experiment tracking, model registry, and versioning for models, datasets, and evaluation configurations; Automate the path from research checkpoints to validated deployments through CI/CD and reproducible workflows; Profile end-to-end model workloads and collaborate with distributed systems, inference, and GPU kernel engineers on deeper performance issues.
Seniority
Mid-level, hands-on IC