CareerPlanGet AI match score →

Site Reliability Engineer - AI Agents

United States💼 Full-time🗓 2026-06-11 → 2026-07-31

Core

Design, build, and operate the infrastructure layer supporting AI agent workflows in production, ensuring reliability, scalability, and observability for agentic systems.

Role type

Senior Site Reliability Engineer (AI Infrastructure & Platform)

Builds

APIs, SDKs, and platform capabilities enabling engineering teams to consume AI infrastructure and agent platform services as a service.

Domain

Financial technology / Applied AI / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Kubernetes, Terraform, AWS, Python, bash/shell scripting, CI/CD pipelines, observability, incident response, containerization, API design, developer experience

Preferred skills

Agent orchestration frameworks (LangGraph, CrewAI), data infrastructure (Airflow, Kafka, Spark), Cloudflare ecosystem, evaluation frameworks

Technologies

Kubernetes, Terraform, AWS, Docker, Python, bash, CI/CD tools, LangGraph, CrewAI, Airflow, Kafka, Spark, Cloudflare

Responsibilities

Design and develop platform services, APIs, and SDKs for self-service infrastructure consumption; manage compute, orchestration, and serving infrastructure for model inference; implement monitoring, alerting, and incident response for AI/ML workloads; define guardrails and failure handling patterns for agentic systems; collaborate to translate agent prototypes into production systems; document architecture and runbooks.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗