Site Reliability Engineer - AI Agents
Core
Design, build, and operate the infrastructure layer supporting AI agent workflows in production, ensuring reliability, scalability, and observability for agentic systems.
Role type
Senior Site Reliability Engineer (AI Infrastructure & Platform)
Builds
APIs, SDKs, and platform capabilities enabling engineering teams to consume AI infrastructure and agent platform services as a service.
Domain
Financial technology / Applied AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, AWS, Python, bash/shell scripting, CI/CD pipelines, observability, incident response, containerization, API design, developer experience
Preferred skills
Agent orchestration frameworks (LangGraph, CrewAI), data infrastructure (Airflow, Kafka, Spark), Cloudflare ecosystem, evaluation frameworks
Technologies
Kubernetes, Terraform, AWS, Docker, Python, bash, CI/CD tools, LangGraph, CrewAI, Airflow, Kafka, Spark, Cloudflare
Responsibilities
Design and develop platform services, APIs, and SDKs for self-service infrastructure consumption; manage compute, orchestration, and serving infrastructure for model inference; implement monitoring, alerting, and incident response for AI/ML workloads; define guardrails and failure handling patterns for agentic systems; collaborate to translate agent prototypes into production systems; document architecture and runbooks.
Seniority
Senior, hands-on IC