AI Benchmarking Lead, Performance Benchmarking Evaluation
Core
Lead performance benchmarking and evaluation for Seller Assistant, Amazon's Gen-AI copilot for sellers, ensuring model reliability and audit quality.
Role type
Senior IC AI Benchmarking Lead (Performance Evaluation)
Builds
Seller Assistant AI model evaluations and audit processes
Domain
E-commerce / Generative AI / Data Annotation
Deliverable
production ML models
Required skills
Natural language data labeling, data annotation, linguistic annotation, process documentation, data analysis, SOP development, stakeholder communication
Preferred skills
SQL querying, Python programming, analytical problem-solving, independent work
Technologies
MS Excel, SQL, Python
Responsibilities
Evaluate audits to increase confidence in evaluation metrics, improve audit reliability through systematic measurement, conduct calibration for quality standards, quality-check audits and provide feedback, drive continuous improvement in audit processes, identify rubric gaps and evaluation ambiguities, surface product issues by validating model failures, modify annotation methods and update SOPs, test new SOPs and tools, structure data collection and analyze results, track and report progress on key metrics, identify operational issues related to process and tooling (via careerplan.io/jobs/10563414-ai-benchmarking-lead-performance-benchmarking-evaluation-at-amazon)