Senior Site Reliability Engineer
Core
Own the reliability, scalability, and operational posture of a multi-cloud infrastructure for an AI-native commerce iPaaS.
Role type
Senior Site Reliability Engineer (Infra-first, AI-assisted development)
Builds
Multi-cloud infrastructure (AWS, GCP, Azure), CI/CD pipelines, observability stacks, incident response workflows, internal tooling
Domain
Cloud Infrastructure, AI/ML Infrastructure, Commerce Technology
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Cloud providers (AWS, GCP, Azure), Observability tooling (Datadog, Grafana), AI-assisted development (Claude Code, Copilot, Cursor), Incident management, SLO/SLI definition, IaC
Preferred skills
MCP or agentic AI infrastructure experience, API gateway infrastructure, Startup/high-growth SaaS experience
Technologies
AWS, GCP, Azure, Terraform, Datadog, Grafana, Kubernetes, Claude Code, Copilot, Cursor
Responsibilities
Own infrastructure across multiple cloud environments; Build and maintain CI/CD pipelines and observability stacks; Define and enforce SLOs/SLIs and lead postmortems; Author and maintain IaC; Write internal tooling and automation using AI-assisted workflows; Partner with engineering on reliability reviews and architecture decisions
Seniority
Senior, hands-on IC