Site Reliability Engineer
Core
Ensure stability and scalability of internal Agentic AI platform through code, automation, observability, and incident response.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Internal Agentic AI platform, LLM proxies, RAG pipelines, and agent gateways
Domain
Financial Services / Artificial Intelligence / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud automation, Python, REST APIs, event-driven architectures, relational databases, NoSQL databases, AWS services, Kubernetes, CI/CD pipelines
Preferred skills
Agentic AI application development, MCP servers, RAG pipelines, DataDog observability
Technologies
Python, AWS, Kubernetes, Github Actions, Postgres, MySQL, DynamoDB, Elastic Search, DataDog
Responsibilities
Define SLOs and error budgets, own observability and incident response, ensure high availability of AI agents and pipelines, support users in dedicated help channel, improve code quality and development lifecycles
Seniority
Senior, hands-on IC