MTS, Forward Deployed Engineer + more at ARTIFICIAL ANALYSIS
Core
Build and run AI benchmarking stacks, evaluations, and harnesses for language models, inference providers, hardware, and speech to support the Intelligence Index.
Role type
Senior IC machine-learning engineer (evaluations & benchmarking)
Builds
Frontier evals, datasets, harnesses, benchmarks, and benchmarking stacks for LLMs, inference, hardware, and speech
Domain
Artificial Intelligence / Machine Learning / Benchmarking (via careerplan.io/jobs/_49928396-mts-forward-deployed-engineer-more-at-artificial-analysis-at-artificial-analysis)
Deliverable
production ML models
Required skills
LLM benchmarking, dataset construction, evaluation harness development, inference provider analysis, hardware benchmarking, speech evaluation (TTS/STT)
Preferred skills
Pre-release model access, working with AI labs, flat structure ownership
Technologies
LLMs, inference providers, GPUs, TPUs, custom silicon, TTS, STT
Responsibilities
Run LLM benchmarking stack and work directly with AI labs, build frontier evals including datasets and harnesses, benchmark speed quality and price across serverless inference providers, benchmark GPUs TPUs and custom silicon, own TTS STT and speech-to-speech evals
Seniority
Senior, hands-on IC
