Senior Site Reliability Engineer I
Core
Own the reliability, observability, and operational excellence of the Foundry data platform, ensuring scalable and resilient data pipelines, Lakehouse infrastructure, and APIs.
Role type
Senior Site Reliability Engineer (Data Platform)
Builds
Foundry data platform services including data pipelines, Lakehouse infrastructure, APIs, and shared capabilities
Domain
Risk analytics, Data Engineering, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux, Kubernetes, Networking, CI/CD, Infrastructure as Code, Containerization, Scripting
Preferred skills
Snowflake, Databricks, ECS/Fargate, Serverless Framework
Technologies
Docker, Kubernetes, GitHub, GitHub Actions, Terraform, CloudFormation, AWS, Azure, Bash, Python
Responsibilities
Design and maintain highly available container-based platforms; Define and operate GitOps-driven delivery workflows; Build and maintain CI/CD pipelines; Monitor and optimize system performance and reliability; Collaborate on best practices for container orchestration; Troubleshoot deployment pipelines and runtime environments; Implement automated scaling and recovery mechanisms; Embed security and compliance into CI/CD workflows.