Member of Technical Staff - MTS
Core
Build and maintain large-scale, distributed, fault-tolerant observability systems (logging, metrics, monitoring) to empower engineering teams in monitoring, troubleshooting, and optimizing web services.
Role type
Senior Site Reliability Engineer (Associate level)
Builds
Cloud-native logging, metrics, and monitoring infrastructure; containerized environments; CI/CD pipelines
Domain
Healthcare technology / Cloud Infrastructure / Observability
Deliverable
infrastructure
Required skills
Linux, Docker, Kubernetes, Python, Bash, scripting, automation, incident response, root cause analysis, capacity planning, CI/CD, infrastructure as code
Preferred skills
Cloud platforms, infrastructure as code tools
Technologies
Docker, Kubernetes, Python, Bash
Responsibilities
Develop and maintain scalable logging, metrics, and monitoring systems; Manage containerized environments; Analyze system performance and reliability metrics; Collaborate with development teams to integrate observability best practices; Automate operational processes; Participate in incident response and root cause analysis; Assist in capacity planning and infrastructure scaling strategies
Seniority
Senior, hands-on IC (Associate level)