CareerPlanGet AI match score →

Software Developer (Agentic Evaluation)

3 Locations💼 Full-time💰 $88,000–$128,700🗓 2026-06-26 → 2026-07-30

Core

Build and rigorously evaluate intelligent agentic systems, including benchmarking AI agents against commercial solvers, to enhance developer productivity and experience.

Role type

Senior IC software developer specializing in agentic AI evaluation

Builds

Multi-agent AI systems for automated test generation, execution, and end-to-end development workflow optimization; MCP-based tooling for IDEs

Domain

Generative AI, Software Engineering, Test Automation

Deliverable

production ML models | product features

Required skills

Python, Large Language Models, AI evaluation methodologies, statistical analysis, experimental design

Preferred skills

QA/Software Engineering background, test automation frameworks, MCP server development, agentic AI frameworks, vision-language models, cloud ML platforms

Technologies

LangGraph, AutoGen, Anthropic Agent SDK, PyTorch, Transformers, scikit-learn, AgentBench, Langfuse, Playwright, Selenium, Pytest, Appium, Azure AI Foundry, AWS

Responsibilities

Develop and orchestrate multi-agent AI systems for automated test generation and execution; Design agentic workflows to coordinate AI agents for test automation across UI, API, and system levels; Build evaluation frameworks and custom benchmarks comparing AI agents against commercial solvers; Evaluate MCP server and tool performance across agentic pipelines

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗