Technical Lead - Site Reliability Engineering
Core
Lead the establishment of SRE foundations for new projects, ensuring operational readiness, monitoring, alerting, and reliability from day one across Markets and Risk Intelligence platforms.
Role type
Senior hands-on IC Technical Lead SRE
Builds
Reliability foundations, observability platforms, monitoring and alerting solutions, resilient system patterns
Domain
Financial markets infrastructure, cloud-native platforms, risk intelligence
Deliverable
production ML models | infrastructure
Required skills
SRE principles (SLOs, error budgets, incident management), Kubernetes, Linux systems administration, Azure (AKS, Container Apps, VNet), Datadog, observability stack design (metrics, logs, traces), cloud cost optimization, AI integration with observability stacks
Preferred skills
AWS, multi-cloud/hybrid environments, Infrastructure as Code (Terraform, CloudFormation), vector databases, RAG architectures, Generative AI/LLM platforms
Technologies
Azure, AKS, Azure Container Apps, VNet, Datadog, Prometheus, Grafana, ELK, OpenTelemetry, Terraform, CloudFormation, Claude, Amazon Bedrock
Responsibilities
Lead establishment of SRE foundations for new projects; Collaborate with Architecture and Engineering to embed reliability, scalability, security, and observability; Define and implement observability standards and tooling; Design monitoring and alerting solutions; Drive reliability improvements through incident reduction and performance tuning; Partner with Security teams on compliance and risk management; Lead handovers from project delivery to BAU operations; Mentor engineers and shape engineering standards
Seniority
Senior, hands-on IC with leadership presence