Software Engineers: Paid Interview on AI Evaluation Tasks
Core
Paid research interview evaluating the quality, accuracy, and realism of programming tasks and technical harnesses used to benchmark AI agents.
Role type
Senior IC software engineer (AI evaluation research)
Builds
AI evaluation benchmarks and testing methodologies
Domain
Artificial Intelligence + Software Engineering
Deliverable
research
Required skills
software engineering, code review, automated testing, technical architecture, evaluation harnesses, critical thinking, analytical skills
Responsibilities
Review and assess the quality and realism of programming tasks for AI agents; Evaluate coding environments and technical evaluation harnesses for correctness; Examine code structures and logic to identify flaws; Assess difficulty and complexity of programming tasks; Provide detailed feedback on task design and evaluation methodology; Discuss technical architecture and testing approaches.
Seniority
Senior, hands-on IC