Systems Integration Specialist Advisor
Core
Senior Site Reliability Engineer supporting cloud operations for microservice-based platforms.
Role type
Senior hands-on IC Site Reliability Engineer
Builds
Cloud infrastructure, automation, and observability for microservice platforms
Domain
Cloud Infrastructure / DevOps / SRE
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
AWS, Azure, Terraform, Atlantis, ArgoCD, Docker, Kubernetes, CI/CD pipelines, incident management, RCA, observability, cloud networking, scripting
Preferred skills
Microservice platforms, Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, Python, Bash, Go, Java, SLI/SLO/SLA, error budgets, capacity planning, disaster recovery testing, mentoring
Technologies
AWS, Azure, Terraform, Atlantis, ArgoCD, Docker, Kubernetes, Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics
Responsibilities
Own and improve reliability of cloud-based services and supporting infrastructure, Participate in on-call rotation and support production systems outside normal business hours, Lead incident response activities including triage, escalation, mitigation, and service restoration, Drive blameless postmortems and ensure corrective actions are tracked to closure, Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis, Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools, Support and improve cloud/container platforms across AWS and Azure, Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems, Build automation to reduce manual effort and improve operational efficiency, Configure and improve monitoring, alerting, logging, diagnostics, and observability, Analyze performance and capacity trends to identify bottlenecks and improve scalability, Troubleshoot complex infrastructure, networking, application runtime, and cloud platform issues, Support disaster recovery planning, validation, and recovery readiness, Create and maintain operational runbooks, support procedures, and engineering documentation, Coach and guide other engineers on SRE best practices, reliability, automation, and operational excellence
Seniority
Senior, hands-on IC