Senior Site Reliability Engineer
Core
Senior Site Reliability Engineer managing cloud infrastructure, Kubernetes clusters, and platform reliability for an AI-powered chat marketing platform.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable, reliable Kubernetes-based infrastructure for Python-based AI services
Domain
Cloud Infrastructure / SRE / AI Platform
Deliverable
production ML models | infrastructure
Required skills
Linux administration, Kubernetes (EKS), Terraform, Helm, CI/CD pipelines, observability (Prometheus, Grafana), cloud security, networking, Nginx
Preferred skills
Ansible, PostgreSQL tuning, PHP production environments, TDD
Technologies
AWS (EC2, ALB/NLB, WAF, IAM, CloudWatch), EKS, Terraform, Helm, GitHub Actions, Prometheus, Grafana, Ansible, Nginx, Python
Responsibilities
Maintain and harden AWS infrastructure; Operate and evolve EKS clusters; Migrate services to Kubernetes; Codify infrastructure with Terraform and Ansible; Build and improve CI/CD pipelines; Own observability efforts; Support OS-level patching and infra hygiene; Partner with engineers on best practices; Create infrastructure documentation
Seniority
Senior, hands-on IC
