[Agentic AI] Systems Evaluation Engineer
Core
Define measurement methodology and benchmarks to quantify customer outcomes, platform performance, and AI workload effectiveness.
Role type
Systems Evaluation Engineer
Builds
Benchmarks and performance reports for AI systems and software platforms
Domain
Agentic AI, Software Engineering, Performance Analytics
Deliverable
production ML models
Required skills
benchmarking methodologies, success metrics definition, customer outcome analysis, executive reporting, software architecture, programming languages
Preferred skills
AI systems evaluation, telemetry analytics, Linux/Windows validation
Technologies
Amazon Web Services, API, Automation, C++, DevOps, Docker, Java, Kubernetes, Microsoft Azure, Python, React.js, SQL
Responsibilities
Define success metrics, Create benchmarking frameworks, Analyze customer outcomes, Produce executive performance reports, Contribute to technology strategy and engineering roadmaps, Supervise external engineering vendors
Seniority
Senior, hands-on IC