Site Reliability Engineer ll
Core
Operate and maintain AWS-hosted MERN applications and large-scale data pipelines to ensure uptime, scalability, and security for healthcare systems.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production healthcare SaaS platform, data ingestion workflows, and automated operational solutions
Domain
Healthcare technology, Cloud Infrastructure, Data Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS core services (Lambda, ECS/EKS, EMR/Glue, EC2, VPC, IAM, CloudWatch), Python (PySpark), Node.js, MERN stack operations, MySQL administration, Terraform/OpenTofu, distributed data orchestration, incident management, HIPAA compliance
Preferred skills
Healthcare/HIPAA experience, crisis management under pressure, stakeholder communication
Technologies
AWS, Lambda, ECS, EKS, EMR, Glue, EC2, VPC, IAM, CloudWatch, PySpark, Node.js, MySQL, Athena, Terraform, OpenTofu, SQS, SNS, RabbitMQ
Responsibilities
Maintain continuous uptime and security of AWS-hosted MERN applications and backend data architectures; Manage and optimize event-driven serverless architectures on AWS Lambda; Monitor and troubleshoot scheduled PySpark data workflows and triage failed jobs; Participate in on-call rotation to mitigate live application outages and data bottlenecks; Engineer automated workflows to eliminate repetitive operational tasks; Build specialized dashboards and alerts for observability; Lead blameless post-mortems for operational failures
Seniority
Senior, hands-on IC