Software Engineer II - Platform & Infrastructure
Core
Build and operate the observability platform (monitoring, metrics, alerting) and developer tooling that powers Abnormal AI's engineering teams and product delivery.
Role type
Senior Platform & Infrastructure Software Engineer
Builds
Observability stack (Prometheus, Chronosphere, Grafana, PagerDuty), data processing pipelines, and internal developer tooling
Domain
SaaS / AI / Observability / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Python, Golang, distributed systems design, observability principles, fault tolerance patterns, incident response, technical design documentation, async communication
Preferred skills
Prometheus (PromQL), Grafana, Chronosphere, PagerDuty, AWS (EKS, ECS, Lambda), Kubernetes, Terraform, CI/CD, Django, gRPC
Technologies
Prometheus, Chronosphere, Grafana, PagerDuty, Airflow, Spark, AWS, Kubernetes, Terraform, Python, Golang, Django, gRPC
Responsibilities
Own the observability stack to ensure real-time system visibility and cost-efficient operations; design and ship platform features and developer tooling to reduce deployment friction; drive SLAs/SLOs for critical shared infrastructure; participate in on-call rotations to triage and resolve production issues; improve system resilience by automating runbooks and refining failure modes; mentor junior engineers and contribute to cross-team technical direction.
Seniority
Senior, hands-on IC with mentorship responsibilities