Site Reliability Engineer – Data Platforms
Core
Establish service governance, operating models, and disaster recovery while performing hands-on AWS infrastructure operations and automation.
Role type
Senior Site Reliability Engineer (Data Platforms)
Builds
Secure, reliable, scalable AWS application infrastructure and operational processes for data platforms
Domain
Biotechnology / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS infrastructure, Infrastructure as Code, CI/CD, Observability, Incident Response, Service Governance, Disaster Recovery Planning, Stakeholder Negotiation, Risk Management
Preferred skills
EKS, Lambda, Glue, EMR, RDS, Security Audit Processes, Agile Delivery
Technologies
AWS, IAM, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, STS, Jira
Responsibilities
Define service handover processes and acceptance criteria; Maintain service catalogues and operating profiles; Lead business continuity and disaster recovery planning; Operate and automate AWS infrastructure; Participate in on-call incident response; Define SLAs and service indicators; Establish security review and audit-readiness processes
Seniority
Senior, hands-on IC