Senior Platform Engineer
Core
Architecting and implementing a modern observability stack (logging, metrics, tracing) and incident response systems for a large-scale cloud environment to reduce toil and improve system visibility.
Role type
Senior Platform Engineer (SRE focus)
Builds
Observability infrastructure, incident management workflows, and self-service troubleshooting tools for hundreds of cloud services.
Domain
Digital Health / Cloud Infrastructure
Deliverable
infrastructure
Required skills
SRE/Platform Engineering, Observability Architecture (Metrics/Logging/Tracing), Agent Configuration & Instrumentation, Systems Thinking, Incident Management, Automation/IaC
Preferred skills
AWS, EKS, Docker, Terraform, Kafka, RDS, Open-source stacks (Prometheus, ELK, Grafana, Jaeger)
Technologies
New Relic, Datadog, OpenTelemetry, Prometheus, ELK, Grafana, Jaeger, Terraform, CloudFormation, Kafka, RDS, EKS, Docker
Responsibilities
Design and deploy scalable logging, metrics, and APM stacks; Establish agent configuration best practices; Create structured logging patterns and cardinality management; Implement distributed tracing; Evaluate and select incident management platforms; Build alerting policies and incident triage workflows; Automate manual incident response tasks; Set SRE practices and mentor the team.
Seniority
Senior, hands-on IC with architectural leadership