Technical Lead - Site Reliability Engineering
Core
Lead the establishment of SRE foundations for new projects, ensuring operational readiness, monitoring, alerting, and reliability across Markets and Risk Intelligence platforms.
Role type
Senior hands-on IC Technical Lead SRE
Builds
Reliability foundations, observability platforms, monitoring and alerting solutions, resilient system patterns
Domain
Financial markets infrastructure, cloud-native platforms, risk intelligence
Deliverable
production ML models | infrastructure
Required skills
SRE principles (SLOs, error budgets, incident management), Kubernetes, Azure (AKS, Container Apps, VNet), Linux administration, observability tooling (Datadog, Prometheus, Grafana, ELK, OpenTelemetry), cloud cost optimization, system design collaboration
Preferred skills
AWS, multi-cloud/hybrid environments, Infrastructure as Code (Terraform, CloudFormation), vector databases, RAG architectures, Generative AI/LLM platforms
Technologies
Azure, AKS, Azure Container Apps, VNet, Datadog, Prometheus, Grafana, ELK, OpenTelemetry, Terraform, CloudFormation, Claude, Amazon Bedrock
Responsibilities
Lead establishment of SRE foundations for new projects including environments, monitoring, and alerting; Collaborate with Architecture and Engineering to embed reliability, scalability, security, and observability into system design; Define and implement observability standards and guidelines; Design monitoring solutions to reduce toil and improve visibility; Drive reliability improvements through incident reduction and performance tuning; Partner with Security to meet compliance and risk expectations; Mentor engineers and shape engineering standards
Seniority
Senior, hands-on IC with leadership presence