Senior Reliability Engineer
Core
Architect and design highly available, fault-tolerant, and scalable infrastructure and systems for a B2B commerce platform serving merchants and retailers across 29 countries.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Highly reliable and scalable cloud infrastructure, automation services (APIs, CLIs), and observability solutions for the BEES B2B platform.
Domain
B2B Commerce / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud platform expertise (AWS, Azure, GCP), Programming (Python, Go, Java), Containerization and orchestration (Docker, Kubernetes), CI/CD practices, Monitoring and observability, Incident response and troubleshooting, GitOps practices, Software engineering and DevOps, API and CLI development.
Preferred skills
Networking concepts (load balancing, CDN, DNS), B2B ecosystem knowledge, Microservices architecture, Open Telemetry (OTEL), Service Mesh (Istio), Cloud certifications, Kubernetes certifications.
Technologies
AWS, Azure, GCP, Python, Go, Java, Docker, Kubernetes, Github Actions, Azure DevOps, New Relic, Dynatrace, Prometheus, Grafana, ELK stack, ArgoCD, Flux, Istio.
Responsibilities
Architect and design highly available, fault-tolerant, and scalable infrastructure; Collaborate with software development teams to ensure reliability and scalability; Develop and implement best practices, guidelines, and standards for SRE processes; Evaluate and recommend suitable technologies for infrastructure, monitoring, and observability; Drive automation efforts through development of APIs, CLIs, and services; Lead incident response and post-incident analysis; Implement robust monitoring, logging, and alerting solutions.
Seniority
Senior, hands-on IC