Senior Machine Learning Engineer - Platform
Role
Seniority - senior
More about the Senior Machine Learning Engineer - Platform role at Zendesk
Job Description
Zendesk is building production AI systems that autonomously resolve customer service tickets across 100,000+ customer accounts. These systems plan multi-step resolutions, execute real actions through live APIs, and close tickets autonomously. We are now scaling beyond individual agent capabilities and investing in the platform foundation that powers reliable, secure, and measurable agentic services across the company. Here, you won’t be building one-off tools or experiments. You’ll be helping create the platform layer that enables teams to build, operate, and scale agentic services across a massive enterprise footprint. Zendesk has the scale, data density, and production urgency to build platform systems that truly matter. You’ll architect and build the core platform services that enable agentic systems at production scale. This includes designing and implementing the infrastructure, control planes, and developer platforms that support planning, execution, memory, evaluation, and multi-agent orchestration. You’ll help define the primitives that product and AI teams build on top of, ensuring these systems are reliable, extensible, observable, and ready for enterprise use. This is ideal if you see agentic AI not as a model problem, but as a platform engineering and distributed systems problem.
What you’ll build
Agentic Service Platform: You will lead the design of the shared platform services that underpin agentic workflows, including runtime orchestration, workflow execution, state management, retries, and fault recovery. You will help shape the abstractions that allow teams to build intelligent services without needing to re-implement the underlying infrastructure every time. Platform APIs and Internal Developer Experience: You will define and evolve the APIs, SDKs, and internal tooling that make the platform easy to adopt. You will partner closely with product, AI, and infrastructure teams to ensure the platform is usable, secure, and consistent across use cases. Memory, Context, and State Systems: You will design platform-level infrastructure for handling memory and context across concurrent sessions, with attention to consistency, isolation, performance, and cost. Reliability, Observability, and Evaluation : You will build and improve the platform’s observability, tracing, and evaluation systems so teams can understand system behaviour, detect regressions, and enforce quality gates in CI/CD. Governance and Safety: You will help design the guardrails, validation layers, and policy enforcement mechanisms required to operate agentic services safely in enterprise environments.
What you’ll bring
Systems Thinker: You have 5+ years of backend experience (Java, Go, or Python). You understand that "Reliability" in an AI system isn't just




