Site Reliability Engineer
Core
Ensuring stability, scalability, and high performance for a SaaS security platform serving global enterprises.
Role type
Site Reliability Engineer (SRE)
Builds
Customer-facing SaaS security platform
Domain
Cybersecurity / SaaS
Deliverable
production ML models
Required skills
Kubernetes, microservices architecture, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana, Golang, Python, autoscaling, version upgrades, cloud service optimization
Preferred skills
Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong
Responsibilities
Support and maintain service quality of the SaaS security platform; Address complex challenges around scalability, reliability, observability, and cost efficiency; Collaborate with Engineering to maintain Helm charts, deployment, monitoring, and CI/CD pipelines; Define and implement service verification strategies in CI/CD to meet SLAs; Optimize CI/CD workflows and developer experience; Participate in 24/7 on-call rotation; Monitor, debug, and optimize production infrastructure on AWS/GCP
Seniority
Mid-level, hands-on IC