CareerPlanGet AI match score →

Principal Software Engineer

United States, Washington, Redmond💼 Full-time🗓 2026-07-20 → 2026-07-31

Core

Architecting and operating AI-driven support workflows, ensuring reliability, security, and customer trust for agentic systems.

Role type

Principal Software Engineer (AI/LLM Systems)

Builds

Production AI agent platforms, support workflows, and citizen developer tools for building AI agents.

Domain

AI Engineering, LLMs, Agentic Systems, Enterprise Support

Deliverable

production ML models | product features | infrastructure

Required skills

C, C++, C#, Java, JavaScript, Python, LLM-powered application development, RAG pipelines, prompt engineering, agent frameworks (Semantic Kernel, LangChain), fine-tuning, non-deterministic system debugging, observability strategy, threat modeling, cross-organization influence, technical mentorship, distributed systems architecture.

Preferred skills

Experience with Agent Harnesses (GitHub Copilot CLI, Coding Agents), Markdown specs/ADRs, YAML configs, zero-touch deployment, feature flagging, industry thought leadership.

Technologies

Semantic Kernel, LangChain, GitHub Copilot CLI, Coding Agents, Markdown, YAML, C, C++, C#, Java, JavaScript, Python

Responsibilities

Own architecture and run-state reliability of AI workflows; lead incident response and retrospectives; define design patterns for adapting AI to business policies; drive security, privacy, and Responsible AI architecture; set observability and monitoring for concurrent AI agents; establish engineering standards and mentor senior/staff engineers; influence technical strategy across organizational boundaries; lead automation and safe deployment practices; define evaluation strategies for non-deterministic systems.

Seniority

Principal, strategy & mentorship

Rewrite
## About the role * Owning the architecture and run-state reliability of AI-driven support workflows * Defining the reliability bar for incident response and live-site health * Serving as a designated responsible individual (DRI) who leads incident retrospectives to identify root causes, owns repair actions, and prevents recurrence across the platform * Driving the design patterns that adapt AI workflows to changing support business policies and operational processes (e.g., SLA calculations, case ownership, escalation models) * Driving customer trust, satisfaction, and sentiment, ensuring AI agents correctly understand intent and guide customers to resolution without degrading experience * Providing thought leadership on the security, privacy, and Responsible AI architecture—including rethinking role-based access control (RBAC), data access, case ownership vs. processing, and data exposure—and assuring visible compliance evidence (e.g., audit trails) across products * Defining the observability, monitoring, and intervention architecture for multiple AI agents operating concurrently at scale * Driving cross-team alignment and establishing scalable engineering standards across engineering, support business, compliance, and platform teams * Negotiating and resolving conflicts around dependency ownership * Driving agreements that align partner teams to the delivery schedule for AI-managed support * Setting the technical vision for and driving delivery of a platform that enables citizen developers to safely build AI agents for support workflows with reduced barrier to entry * Raising the engineering bar across the team by mentoring senior and staff engineers, leading design and code reviews, and establishing best practices for building and evaluating production agentic systems * Influencing technical strategy and roadmaps across organizational boundaries, building alignment among partner teams and senior leaders behind a shared architectural vision for AI-managed support * Leading the application of automation across production and deployment for complex products, targeting zero-touch deployment, and driving safe change-deployment best practices (e.g., correct flighting, rollback plans) to minimize customer impact * Leading experimentation using feature flags/flighting to measure the impact of changes, partnering with Data Science and product managers to define the success and guardrail metrics that drive customer value * Partnering with and guiding internal and external stakeholders to anticipate, determine, and confirm customer/user requirements and their feasibility, and advocating for the security and privacy needs of the customers using the platform ## Requirements * Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience * Experience building production LLM-powered applications, RAG pipelines, prompt engineering, agent frameworks (Semantic Kernel, LangChain), and fine-tuning, with strong judgment on evaluation, latency, and cost trade-offs at scale * Experience in AI-native development working within Agent Harnesses (GitHub Copilot CLI, Coding Agents), authoring Markdown specs/ADRs and YAML configs as Agent-consumable inputs, orchestrating multi-step Agentic workflows across the SDLC, and setting the standard for reviewing Agent-generated code and PRs with production-grade rigor * Experience shipping and operating agent-based systems in production at scale, across multiple teams or products, including evals, observability, and debugging of non-deterministic behavior * Experience defining the evals and observability strategy for non-deterministic systems and driving its adoption across teams * Contributions to the safety posture of AI systems, including prompt-injection defences, threat modeling, and audit trails * Experience owning and driving the architecture, design, and delivery of entire products or solutions that are deeply complex and often ambiguous, with measurable impact well beyond a single team * Cross-organization and company-wide influence: you build consensus among senior leaders and Principal/Partner engineers, resolve deep ambiguity, and drive large, complex initiatives to outcomes * Technical leadership: you set technical direction across multiple teams or an organization, mentor senior and Principal engineers, and establish engineering standards and best practices that others adopt * Experience architecting and operating company-scale, mission-critical distributed systems with stringent reliability, security, privacy, and compliance requirements * Industry or company-wide thought leadership in AI engineering—for example, driving adoption of emerging technologies, representing engineering to executives and external audiences, or contributions such as patents, publications, or widely adopted internal frameworks
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Microsoft ↗