CareerPlanGet AI match score →
💼 Full-time🗓 2026-06-25

Core

Design rigorous software engineering challenges and benchmark frontier AI models against real-world codebases to evaluate model capabilities.

Role type

AI Research Engineer (Software Engineering Benchmarks)

Builds

Benchmark datasets and evaluation pipelines for frontier AI models

Domain

Artificial Intelligence / Software Engineering

Deliverable

production ML models

Required skills

Software Engineering, Algorithms, Data Structures, Docker, Containerization, Codebase Analysis, Model Evaluation

Preferred skills

AI Research, Frontier Model Evaluation

Technologies

Docker

Responsibilities

Create complex software engineering challenges across diverse domains, Curate high-fidelity training and benchmarking data using real-world codebases, Build containerized, reproducible training pipelines, Systematically evaluate and benchmark frontier models against designed problems

Rewrite
## About the Role In this role, you will be at the intersection of high-level research and deep infrastructure. You will design rigorous challenges that push frontier models to their limits, ensuring that the next generation of AI is capable of handling real-world software complexity. ## Key Responsibilities - Problem Design: Create complex software engineering challenges across diverse domains, including algorithms, computer science, and embedded systems. - Data Generation: Curate high-fidelity training and benchmarking data using real-world codebases to evaluate the world's most capable models. - Infrastructure: Leverage Docker-heavy environments to build containerized, reproducible training pipelines. - Model Benchmarking: Systematically evaluate and benchmark frontier models against the hard problems designed by our research team. ## Requirements - Technical Expertise: Strong background in Software Engineering, Algorithms, and Data Structures. - Infrastructure Skills: Heavy experience with Docker and containerization for building scalable research infrastructure. - Problem-Solving: Ability to deconstruct complex, real-world codebases into structured training data. - Research Mindset: Experience in or a strong passion for AI research and evaluating frontier models. ## Why TutorDock AI? You will join a focused research team dedicated to solving the most difficult problems in AI for software engineering. If you are passionate about building the benchmarks that will define the future of AI capabilities, we want to hear from you.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗