Senior Reliability Engineer
Core
Senior Service Reliability Engineer ensuring excellence in Fitch Solutions services with a focus on new AI development and mission-critical cloud infrastructure.
Role type
Senior Service Reliability Engineer (SRE)
Builds
Cloud infrastructure, platform engineering, and AI-enabled operational services for Fitch Solutions development squads.
Domain
Financial services technology, Cloud Infrastructure, AI Operations
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, AWS, Azure, Docker, Linux, Windows, .NET, Java Spring Boot, Python, PowerShell, Bash, CI/CD (GitHub Actions), Observability (Datadog), Cloud Security (IAM, SCPs, OPA), AI/ML workloads (SageMaker, AWS Bedrock, MCP)
Preferred skills
Agentic AI for operations, Policy-as-code (OPA), CSPM tools (Wiz), Agile delivery
Technologies
Kubernetes, AWS, Azure, Docker, GitHub Actions, Datadog, AWS Bedrock, SageMaker, Model Context Protocol (MCP), Wiz, OPA
Responsibilities
Lead delivery of reliable, scalable, mission-critical services; guide squads on Kubernetes and modern deployment patterns; mentor associate engineers; partner with development squads to design service builds and DevOps tooling; architect and govern GitHub Actions CI/CD with quality gates and AI-assisted redeploy checks; own observability including SLIs/SLOs, dashboards, and alerting; champion AI-enabled operations for log analysis and incident triage; define and enforce cloud guardrails and security controls; influence cross-functional roadmaps and lead complex release planning.
Seniority
Senior, hands-on IC with mentorship responsibilities