Platform Engineer, Observability Platforms
Core
Design, build, and operate scalable observability platform capabilities including monitoring, alerting, tracing, and event management to enable reliable service resilience across the company.
Role type
Senior Platform Engineer (Observability & Infrastructure)
Builds
Scalable observability platform services, automated CI/CD workflows, and cloud-native infrastructure.
Domain
Cloud Engineering, Observability, DevOps
Deliverable
production ML models | product features | infrastructure
Required skills
System Design, CI/CD Pipeline Development, Cloud Engineering (AWS), API Development (Python), Observability & Monitoring (OpenTelemetry, Prometheus, Grafana), AIOps (LLMs), Agile/Scrum
Preferred skills
Container Orchestration (Docker, Kubernetes), DevSecOps, Bash Scripting, SDLC Processes
Technologies
AWS (EC2, S3, Lambda, ECS/EKS, RDS, IAM, VPC, CloudWatch), GitLab CI, GitHub Actions, Terraform, CloudFormation, Python (Flask/FastAPI), OpenTelemetry, Prometheus, Grafana, Big Panda, Logic Monitor, Docker, Kubernetes, ELK Stack, Datadog
Responsibilities
Design and automate infrastructure and CI/CD workflows using IaC; Lead engagements with product teams to define custom observability solutions; Drive continuous improvement through AIOps and telemetry insights to reduce alert noise.
Seniority
Senior, hands-on IC with advisory responsibilities