CareerPlanGet AI match score →

Senior AI Product Engineer, Backend

💼 Full-time🗓 2026-06-25

Core

Building scalable distributed backend services and APIs that power an ML observability platform for engineering teams.

Role type

Senior IC backend engineer (distributed systems & ML infrastructure)

Builds

High-volume analytics systems, real-time evaluation infrastructure, and collection tools for ML/LLM pipelines.

Domain

Cloud infrastructure, distributed systems, and ML observability

Deliverable

production ML models | product features | infrastructure

Required skills

Go, Python, TypeScript/Node, Java, high-performance backend systems, OLAP database architecture, distributed message queues, container orchestration (Kubernetes), public clouds (AWS/GCP/Azure)

Preferred skills

Kafka, stream processing, system observability tooling (Prometheus), dimensionality reduction algorithms, LLM development

Technologies

Go, Python, TypeScript, Java, Kubernetes, AWS, GCP, Azure, Kafka, Prometheus

Responsibilities

Write maintainable, scalable backend code; design and build APIs for ML/LLM workflows; prototype and optimize backend services; extend open source OLAP databases and message queue frameworks; develop collection tools for monitoring ML/LLM pipelines; research and implement visualization & dimensionality reduction algorithms; collaborate with product and customer engineering teams; contribute to in-house AI Agents.

Seniority

Senior, hands-on IC

Rewrite
## About the Role Our Backend Engineering team builds all of the highly scalable distributed services that power Arize's ML observability platform. While Go is our primary language for these distributed systems, the team also maintains services and tools written in Python, Java, and TypeScript. The expectation and scope of every individual on this team is high, whether it's finding the most efficient way to compute model evaluation metrics across billions of data points, designing the next generation of our OLAP database architecture, or researching and implementing the latest dimensionality reduction techniques – you will never lack a technical challenge. You will be a part of the core team that drives product innovation at Arize. You will be challenged with understanding how some of the most impactful engineering teams are developing AI and LLM-powered applications, and how to build the right tools to enable them to do their best work. Our product solutions range from clean APIs that magically instrument applications, interactive playgrounds for prompt engineering and agent development, or scaling up real-time evaluation infrastructure to handle millions of annotations per second. ## Responsibilities - Write maintainable, scalable, and performant backend code primarily in Go, Java, and Python, with opportunities to work in TypeScript. - Build high-volume and highly available analytics systems. - Design and build APIs specific to our customers' Machine Learning and LLM workflows. - Prototype, optimize, and maintain scalable backend services that power the Arize core platform. - Extend, and contribute back to, open source OLAP databases and distributed message queue frameworks. - Develop and integrate collection tools for robust monitoring of ML and LLM pipelines. - Research and implement cutting-edge visualization & dimensionality reduction algorithms in a distributed environment. - Collaborate with our product, design, and directly with customer engineering teams to enhance and expand our product offerings. - Contribute to the build our own in-house AI Agents ## Requirements - 5+ years of experience working with high-performance backend systems. - Strong experience writing Go, Python, TypeScript/Node, Java, or similar server programming languages. - Enthusiasm and interest in the AI and LLM ecosystem, with a desire to learn and stay updated on emerging technologies. - Previous work building and operating highly complex SaaS platforms/systems. - Knowledge of working with public clouds & container orchestration - AWS, GCP, Azure, Kubernetes, etc. ## Nice to Have - Experience with distributed stream processing - Kafka, Gazette, or similar. - Experience with OLAP systems. - Familiarity with system observability tooling like Prometheus. - Working knowledge of Machine Learning and/or Data Science. - First-hand experience working with large language models (LLMs) or developing AI products.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗