Sr. Specialist - Platform Operations (AI & Agentic Systems)
Core
Manage, scale, and optimize production-grade Kubernetes clusters and AI/agentic systems to ensure reliable platform operations for global markets and clients.
Role type
Senior Infrastructure Engineer (Kubernetes & AI Operations)
Builds
Production-grade Kubernetes clusters, automated deployment pipelines, and operationalized AI/agentic services.
Domain
Financial Services / Cloud Infrastructure / AI Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes administration, GitOps (ArgoCD), Infrastructure as Code (Terraform/OpenTofu/Pulumi), Observability (Prometheus/Grafana/ELK), Linux internals, Cloud architecture (AWS), Python/Go/Bash scripting
Preferred skills
Internal Developer Platforms (Backstage), Service meshes (Istio/Linkerd), MLOps, AI observability and cost optimization
Technologies
Kubernetes, ArgoCD, Terraform, OpenTofu, Pulumi, Prometheus, Grafana, ELK, OpenSearch, AWS, GitHub Actions, GitLab CI, Jenkins, Docker, containerd, Istio, Linkerd, Backstage
Responsibilities
Manage and optimize multi-cloud Kubernetes clusters; Design and maintain GitOps pipelines; Implement monitoring and alerting systems; Conduct post-mortems and automate operational toil; Collaborate on developer experience improvements; Deploy and monitor AI/agentic workflows; Implement automation for system testing and deployment.
Seniority
Senior, hands-on IC