Senior Site Reliability Engineer
Core
Establish and evolve SRE best practices, define observability strategy, and design software-driven infrastructure solutions to ensure platform resilience and scalability for Blink Health's prescription access products.
Role type
Senior Site Reliability Engineer (Platform & Infrastructure)
Builds
Production-grade cloud infrastructure, observability systems, and automation tooling for BlinkRx and Quick Save platforms.
Domain
Healthcare technology, Cloud Infrastructure, Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux systems fundamentals, Kubernetes and container orchestration, Python and Go programming, Infrastructure as Code (Terraform/Pulumi), Networking (TCP/IP, DNS, Load Balancing), Observability (metrics, logging, tracing), Incident response and postmortems, Agile methodologies
Preferred skills
Service mesh concepts, Microservices architecture, React troubleshooting, Security scanning and secrets management, Cost optimization
Technologies
AWS, GCP, Azure, Kubernetes, EKS, Helm, Terraform, Pulumi, CloudFormation, Ansible, Python, Go, Bash, React
Responsibilities
Establish and evolve SRE best practices including error budgets and operational readiness; Define and drive observability strategy including SLIs/SLOs and alerting quality; Design and implement software-driven solutions to automate manual processes; Act as a technical leader influencing decision-making across core cloud infrastructure; Take ownership of large, ambiguous initiatives from concept to delivery; Proactively identify systemic risks and lead platform upgrades; Partner with engineering teams to improve developer workflows and operational maturity; Provide technical mentorship, architecture guidance, and code reviews; Lead incident response practices and post-incident learning.
Seniority
Senior, hands-on IC with strategic influence