Site Reliability Engineer
Core
Design, implement, and manage on-premise Kubernetes infrastructure and MLOps platforms for a defence AI company.
Role type
Site Reliability Engineer (Infrastructure & MLOps)
Builds
Cloud-native infrastructure platforms, observability frameworks, secure multi-tenant Kubernetes clusters, and MLOps pipelines.
Domain
Defence / AI / Cloud-Native Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (production operations, operators, service mesh), Observability (Grafana, Prometheus, OpenTelemetry), Scripting (Python, Go, Rust, Bash), Networking (security, zero-trust), Infrastructure as Code (Terraform, Ansible), MLOps platforms (Kubeflow, MLflow), Linux system administration.
Preferred skills
GitOps workflows, CI/CD automation, Policy-as-code (OPA/Gatekeeper), Container security (Falco).
Responsibilities
Design and build cloud-native infrastructure platforms on-premises; create robust observability frameworks; architect and implement secure, multi-tenant Kubernetes clusters; build and maintain MLOps platforms; collaborate with Security teams on supply chain security and runtime protection.
Seniority
Mid-to-Senior, hands-on IC