Senior Site Reliability Engineer
Core
Own the availability, performance, and operability of HTTP APIs and an async compute platform for a creative AI node graph editor.
Role type
Senior Site Reliability Engineer
Builds
HTTP APIs for plugin management and an async compute platform for running AI graphs on the cloud
Domain
Cloud infrastructure, AI creative tools, SRE
Deliverable
production ML models | infrastructure
Required skills
Kubernetes production expertise, Docker containerization, Node.js/TypeScript, Postgres/Redis, Terraform, AWS, observability tooling, CI/CD automation, database backup and disaster recovery
Preferred skills
HTTP API security, blameless postmortem culture
Technologies
Kubernetes, Docker, Node.js, TypeScript, Postgres, Redis, AWS, Aurora, CircleCI, Terraform
Responsibilities
Define and enforce SLOs, SLIs, and error budgets; build and maintain observability (metrics, logging, tracing); lead incident response and postmortems; improve reliability of async job scheduling; maintain CI/CD systems; design cloud infrastructure automation; reduce operational toil through tooling
Seniority
Senior, hands-on IC