Site Reliability Engineer (SRE) - DevOps
Core
Build and maintain reliable, scalable AWS infrastructure and CI/CD pipelines for enterprise SaaS applications, ensuring high availability and operational efficiency.
Role type
Senior IC Site Reliability Engineer (DevOps)
Builds
Production-ready cloud infrastructure, automated deployment pipelines, and observability dashboards for SaaS platforms.
Domain
Cloud Infrastructure / DevOps / SaaS
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS (EC2, VPC, IAM, RDS, ECS, S3), Terraform, CI/CD pipeline design (GitHub Actions, Jenkins, GitLab CI/CD), Datadog, Docker, Kubernetes, Python, Bash scripting, incident response, root cause analysis, on-call rotation management.
Preferred skills
SRE best practices (SLIs, SLOs, error budgets), internal tooling development, multi-region AWS architecture, FinOps/cost optimization, high-growth B2B SaaS experience.
Technologies
AWS, Terraform, Datadog, GitHub Actions, Jenkins, GitLab CI/CD, Docker, Kubernetes, Python, Bash
Responsibilities
Participate in bi-weekly on-call rotations to monitor production systems and respond to incidents; perform root cause analysis and implement permanent fixes; build and manage AWS infrastructure using Terraform; design, implement, and maintain CI/CD pipelines; automate repetitive operational tasks and improve internal tooling; build and maintain monitoring, dashboards, alerting, and observability using Datadog; partner with software engineering teams to improve operational readiness before production releases; remediate infrastructure-related vulnerabilities; develop and maintain operational documentation, standards, and runbooks.
Seniority
Senior, hands-on IC