Software Engineer, Agent Evaluation and Quality
Core
Building measurement, evaluation, and feedback-loop infrastructure to ensure the reliability and quality of Cursor's core AI coding agent.
Role type
Senior IC software engineer (AI evaluation systems)
Builds
AI evaluation systems, feedback pipelines, analysis tooling, and quality guardrails for an automated coding agent.
Domain
AI/ML engineering, software quality assurance, automated coding tools
Deliverable
production ML models | product features
Required skills
AI evaluation system design, data pipeline engineering, statistical analysis, software engineering fundamentals, debugging complex systems
Preferred skills
Experience with experimentation or ranking systems, strong data acumen, knowledge of emerging AI research trends
Technologies
(Not explicitly stated)
Responsibilities
Designing and building best-in-class AI evaluation systems with curated datasets and scorers; developing feedback loops from real usage data; creating analysis tooling for debugging agent behavior; improving reliability and guardrails by making quality measurable and operational
Seniority
Senior, hands-on IC