Software Developer (Agentic Evaluation)
Core
Build and rigorously evaluate intelligent agentic systems, including benchmarking AI agents against commercial solvers, to enhance developer productivity and experience.
Role type
Senior IC software developer specializing in agentic AI evaluation
Builds
Multi-agent AI systems for automated test generation, execution, and end-to-end development workflow optimization; MCP-based tooling for IDEs
Domain
Generative AI, Software Engineering, Test Automation
Deliverable
production ML models | product features
Required skills
Python, Large Language Models, AI evaluation methodologies, statistical analysis, experimental design
Preferred skills
QA/Software Engineering background, test automation frameworks, MCP server development, agentic AI frameworks, vision-language models, cloud ML platforms
Technologies
LangGraph, AutoGen, Anthropic Agent SDK, PyTorch, Transformers, scikit-learn, AgentBench, Langfuse, Playwright, Selenium, Pytest, Appium, Azure AI Foundry, AWS
Responsibilities
Develop and orchestrate multi-agent AI systems for automated test generation and execution; Design agentic workflows to coordinate AI agents for test automation across UI, API, and system levels; Build evaluation frameworks and custom benchmarks comparing AI agents against commercial solvers; Evaluate MCP server and tool performance across agentic pipelines
Seniority
Senior, hands-on IC