Member of Technical Staff, Coding Research
Core
Design evaluation frameworks, benchmarks, and data systems to measure and improve frontier coding agents.
Role type
Senior IC research engineer (coding agent evaluation)
Builds
Evaluation pipelines, datasets, and tooling for coding model assessment
Domain
AI research / Software Engineering
Deliverable
production ML models
Required skills
Python, C++, software engineering, machine learning, AI research, model evaluation, benchmark design, data generation, tooling development, analytical skills
Preferred skills
frontier AI systems, coding agents, reinforcement learning, agentic workflows, post-training methodologies, technical leadership
Technologies
Python, C++, large language models, reinforcement learning
Responsibilities
Design evaluation frameworks and scoring methodologies for coding agents; Lead research initiatives to measure model performance; Develop high-quality datasets and golden examples; Analyze model behavior and failure modes; Build tooling for large-scale experimentation and evaluation pipelines; Partner with researchers and engineers on experiments; Contribute to technical reports and benchmark studies
