Senior Reliability Engineer
Core
Architect and design highly available, fault-tolerant, and scalable infrastructure and systems for a global B2B commerce platform.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Highly available, fault-tolerant, and scalable infrastructure and systems for merchants and retailers across 29 countries
Domain
B2B commerce, Cloud Infrastructure, SRE
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud platform expertise (AWS, Azure, GCP), Infrastructure-as-Code (Terraform), Programming (Python, Go, Java), Containerization and orchestration (Docker, Kubernetes), Monitoring and observability (New Relic, Dynatrace, Prometheus, Grafana, ELK), Incident response and root cause analysis, Service-level objectives (SLOs) and error budget management, Automation development
Preferred skills
Advanced networking concepts (load balancing, CDN, DNS), B2B ecosystem knowledge, Cloud platform certifications
Technologies
AWS, Azure, GCP, Terraform, Python, Go, Java, Docker, Kubernetes, New Relic, Dynatrace, Prometheus, Grafana, ELK stack
Responsibilities
Architect and design highly available, fault-tolerant, and scalable infrastructure; Collaborate with software development teams to influence design decisions; Develop and implement best practices for SRE processes and automation; Define and enforce SLOs and error budgets; Lead incident response and post-incident analysis; Implement robust monitoring, logging, and alerting solutions; Evaluate and recommend suitable technologies for infrastructure and observability
Seniority
Senior, hands-on IC