Staff Platform Engineer (MANTL)
Core
Lead investigation, troubleshooting, and resolution of complex reliability issues within the MANTL platform application codebase, including microservice communication failures and correctness problems under failure conditions.
Role type
Staff Platform Engineer (Reliability focus)
Builds
The MANTL platform (digital sales and service platform for U.S. banks and credit unions)
Domain
Fintech / Banking / SaaS
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
TypeScript, Node.js, Kubernetes, CI/CD pipeline design, distributed tracing, observability (APM), microservice troubleshooting, performance optimization, fault-injection testing, SLO/SLI definition
Preferred skills
Kafka, OpenTelemetry, regulated environment experience, Infrastructure-as-Code (Terraform)
Technologies
TypeScript, Node.js, Kubernetes, Datadog, GitHub Actions, Kafka, OpenTelemetry, Terraform
Responsibilities
Lead investigation and resolution of complex reliability issues; Set direction for failure mode identification and resilience strategies; Own design and evolution of monitoring, dashboards, and alerting strategy; Establish standards and lead implementation of distributed tracing; Lead diagnosis and remediation of application performance issues; Set direction for platform hardening practices and resilience testing; Own and improve CI/CD build pipeline architecture; Guide deployment and troubleshooting of container-native workloads; Define and drive execution of a roadmap for known reliability risks; Establish and maintain documentation and runbook standards; Serve as senior escalation point for complex platform reliability issues; Define reliability targets (SLOs/SLIs) and advise leadership on risk tradeoffs; Provide technical mentorship to other Platform Engineers
Seniority
Staff, hands-on IC with strategic scope