Principal Software Engineer
Core
Sets technical direction for the Copilot offline evaluation platform, defining scenarios, metrics, and scorecards to gate AI agent releases; leads architecture for an AI-first, agentic-first evaluation platform where autonomous agents are first-class operators.
Role type
Principal Software Engineer (AI Evaluation & Agentic Systems)
Builds
AI-first, agentic-first evaluation platform with autonomous agents as operators, including services, pipelines, tooling, observability, and guardrails.
Domain
AI/ML, Agentic Systems, Large-scale Distributed Systems, Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Large-scale distributed systems design, Reliability engineering, AI/ML evaluation and benchmarking, Autonomous agent system design, API and tooling design for agents, Observability and safety mechanisms, Graceful degradation and intelligent retry patterns, Dependency-aware gating, Modernization of evaluation runtimes, Mentorship of engineering teams.
Preferred skills
Experience with .Net, Java, JavaScript, Rust, or Python, Deep experience with LLM-based grading and metric design, Experience with simulation, scraping, and scoring capabilities for agent behavior.
Technologies
.Net, Java, JavaScript, Rust, Python
Responsibilities
Define end-to-end evaluation for Copilot's agentic experiences (multi-turn trajectories, tool selection, task outcomes, UX, voice), design simulation and scoring capabilities, mentor engineers to produce extensible systems, drive modernization of the evaluation runtime to scale with demand, stay at the frontier of agentic systems and share knowledge across the organization.
Seniority
Principal, strategy & mentorship