CareerPlanSign in

MTS, Forward Deployed Engineer + more at ARTIFICIAL ANALYSIS

San Francisco, CA (ON-SITE, 5 days) or Melbourne, Australia (ON-SITE, 3 days)💼 Full-time🗓 2026-10-01 → 2026-10-03

Core

Build and run AI benchmarking stacks, evaluations, and harnesses for language models, inference providers, hardware, and speech to support the Intelligence Index.

Role type

Senior IC machine-learning engineer (evaluations & benchmarking)

Builds

Frontier evals, datasets, harnesses, benchmarks, and benchmarking stacks for LLMs, inference, hardware, and speech

Deliverable

production ML models

Required skills

LLM benchmarking, dataset construction, evaluation harness development, inference provider analysis, hardware benchmarking, speech evaluation (TTS/STT)

Preferred skills

Pre-release model access, working with AI labs, flat structure ownership

Technologies

LLMs, inference providers, GPUs, TPUs, custom silicon, TTS, STT

Responsibilities

Run LLM benchmarking stack and work directly with AI labs, build frontier evals including datasets and harnesses, benchmark speed quality and price across serverless inference providers, benchmark GPUs TPUs and custom silicon, own TTS STT and speech-to-speech evals

Seniority

Senior, hands-on IC