Site Reliability Engineer
Core
Define and execute the reliability strategy for Yuno's AI agent infrastructure, ensuring the platform powering global payments remains observable, resilient, and scalable.
Role type
Staff Site Reliability Engineer (Infrastructure & Reliability Strategy)
Builds
AI agent provisioning and deployment platform on AWS supporting payments across 190+ countries
Domain
Fintech / Payments / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Event-driven architecture, AWS (EC2, VPC, IAM, S3, RDS), Infrastructure as Code (Terraform/Pulumi), Kubernetes, Observability (Datadog), Chaos engineering, Distributed systems debugging, SQL/NoSQL
Preferred skills
AI/MLOps infrastructure, Multi-tenant container platforms, Data pipelines orchestration
Technologies
Kafka, NATS, RabbitMQ, Terraform, Pulumi, Kubernetes, Docker, Datadog, AWS, PostgreSQL, MongoDB, Redis, Go, Python
Responsibilities
Define SLO culture and error-budget policies, drive architectural decisions for the messaging layer, own cloud infrastructure automation, build monitoring and tracing systems, lead incident response and postmortems, mentor engineering teams
Seniority
Staff, hands-on IC with strategic scope