Associate Site Reliability and Forward Deployed Engineer
Core
Early-career individual contributor supporting production operations, incident response, and post-launch reliability for the Abacus Insights healthcare data platform.
Role type
Associate Site Reliability and Forward Deployed Engineer
Builds
Production reliability, operational automation, and incident resolution for AWS and Databricks-based healthcare data systems
Domain
Healthcare data infrastructure, Cloud Operations, Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Python programming, Linux command-line tools, SQL troubleshooting, AWS cloud fundamentals, incident response, root cause analysis, monitoring and alerting, CI/CD pipelines, Infrastructure as Code concepts
Preferred skills
Databricks, Spark, Snowflake, Kubernetes, EKS, EMR, Lambda, Terraform, ticketing and incident management tools
Technologies
AWS, Databricks, Python, Linux, SQL, Kubernetes, Terraform
Responsibilities
Participate in production incident triage, mitigation, recovery, and follow-up activities; Contribute evidence and findings to root cause analyses (RCAs); Investigate and resolve well-scoped production defects and field-reported issues; Collaborate with Engineering, Data, and Customer Success teams during customer-impacting incidents; Maintain and improve runbooks, troubleshooting guides, monitoring, and alerting; Write and maintain Python scripts to automate operational workflows; Support customer deployments and production troubleshooting; Participate in on-call or shift-based support rotation
Seniority
Associate, early-career IC with mentorship