CareerPlanSign in

AI Benchmarking Lead, Performance Benchmarking Evaluation

Hyderabad, Telangana, India💼 Full-time🗓 2026-09-29 → 2026-10-07

Core

Lead performance benchmarking and evaluation for Seller Assistant, Amazon's Gen-AI copilot for sellers, ensuring model reliability and audit quality.

Role type

Senior IC AI Benchmarking Lead (Performance Evaluation)

Builds

Seller Assistant AI model evaluations and audit processes

Domain

E-commerce / Generative AI / Data Annotation

Deliverable

production ML models

Required skills

Natural language data labeling, data annotation, linguistic annotation, process documentation, data analysis, SOP development, stakeholder communication

Preferred skills

SQL querying, Python programming, analytical problem-solving, independent work

Technologies

MS Excel, SQL, Python

Responsibilities

Evaluate audits to increase confidence in evaluation metrics, improve audit reliability through systematic measurement, conduct calibration for quality standards, quality-check audits and provide feedback, drive continuous improvement in audit processes, identify rubric gaps and evaluation ambiguities, surface product issues by validating model failures, modify annotation methods and update SOPs, test new SOPs and tools, structure data collection and analyze results, track and report progress on key metrics, identify operational issues related to process and tooling (via careerplan.io/jobs/10563414-ai-benchmarking-lead-performance-benchmarking-evaluation-at-amazon)