Site Reliability Engineer - India
Core
Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP. (via careerplan.io/jobs/9bfc9bfb-15f0-4984-85dc-d1adc90a415c-site-reliability-engineer-india-at-jumpcloud)
Role type
Senior IC Site Reliability Engineer
Builds
Cloud-native infrastructure, observability frameworks, and automation tools for a unified IT management platform
Domain
Cloud infrastructure (AWS/GCP), SRE, and IT management
Required skills
Python, Go, Kubernetes, Terraform, AWS, GCP, Datadog, GitOps, Disaster Recovery, Service Meshes
Preferred skills
CI/CD tools, Chaos Engineering, Secrets Management, DevSecOps
Technologies
AWS, GCP, Kubernetes, EKS, Terraform, Datadog, Argo CD, Kargo, Python, Go, Istio, HAProxy, NGINX, GitHub Actions, GitLab Pipelines, HashiCorp Vault
Responsibilities
Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP; Operationalize SLIs, SLOs, and error budgets in direct partnership with core application teams; Build and refine end-to-end observability across microservices and cloud infrastructure using tools like Datadog; Implement actionable monitoring across Golden Signals (Latency, Traffic, Errors, Saturation) to optimize detection (MTTD) and minimize alert fatigue; Participate in on-call rotations, incident response, and blameless post-incident reviews to drive continuous systemic improvements; Manage and operationalize production Kubernetes (EKS) clusters utilizing GitOps delivery workflows (Argo CD, Kargo); Provision and secure multi-cloud infrastructure using modular Terraform (Infrastructure-as-Code); Develop and maintain Disaster Recovery (DR) dashboards, runbooks, multi-region failover automation, and validation tests to ensure alignment with defined RTO and RPO targets; Eliminate operational toil by writing production-grade Python or Go scripts and automation tools; Leverage AI-assisted development tools (Cursor, Claude Code, GitHub Copilot) to accelerate scripting, runbook generation, and incident triage.