Site Reliability Engineer / Anwendungsbetrieb (m/w/d) - KI & AWS Cloud
Core
Operate and monitor AI-based solutions and server infrastructure on AWS, ensuring stability for a high-volume securities transaction platform.
Role type
Senior Site Reliability Engineer (AppOps)
Builds
Event-Driven Microservices architecture on Spring Boot and Kafka processing millions of securities transactions
Domain
Financial Services / Cloud Infrastructure / AI Operations
Deliverable
production ML models | infrastructure
Required skills
AWS Cloud operations, Kubernetes, Prometheus Stack, Microservices architecture, DevOps automation, Incident management, Linux/OS administration
Preferred skills
AWS Bedrock, regulated industry experience, fluent German (C1), English (B2)
Technologies
AWS, Spring Boot, Kafka, Kubernetes, Prometheus
Responsibilities
Operate and monitor AI solutions on AWS, manage server operations and updates, automate deployment processes, handle monitoring and incident response, coordinate releases with platform teams, provide 24/7 on-call coverage
Seniority
Senior, hands-on IC
