Senior Site Reliability Engineer
Core
Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications while owning infrastructure for AI tooling and CI/CD systems.
Role type
Senior Site Reliability Engineer (Infrastructure & AI Tooling)
Builds
Kubernetes clusters, MCP servers for AI access, CI/CD pipelines, streaming/analytics infrastructure
Domain
Cybersecurity SaaS, Cloud Infrastructure, AI/LLM Operations
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals, CI/CD pipeline design, Infrastructure as Code (Terraform/Helm), Python/Go/Bash, observability tooling, Kafka/Flink/ClickHouse, secure AI access patterns
Preferred skills
Multi-cluster Kubernetes, chaos engineering, security scanning/policy-as-code, LLM observability tooling, open-source contributions
Technologies
Kubernetes, EKS/GKE/AKS, Terraform, Helm, Argo CD, GitHub Actions/Jenkins/GitLab CI, Prometheus, Grafana, Datadog, OpenTelemetry, Kafka, Flink, ClickHouse, Langsmith, Langfuse
Responsibilities
Design and scale Kubernetes infrastructure for secure, multi-tenant applications; Build and operate AI tooling infrastructure including MCP servers; Optimize and maintain CI/CD pipelines; Implement progressive delivery strategies; Advance Infrastructure as Code patterns; Operate streaming and analytics infrastructure; Build automated testing into CI/CD; Improve system observability with SLOs and alerts; Lead incident response and postmortems; Mentor engineers on infrastructure practices
