Site Reliability Engineer
Core
Champion performance and stability for the core application ecosystem, bridging code and infrastructure to ensure high-frequency data ingestion pipelines and customer-facing applications meet strict benchmarks.
Role type
Site Reliability Engineer (Java/C#)
Builds
High-availability observability pipelines, integration tests for third-party social APIs, and Grafana dashboards for JVM/.NET environments.
Domain
Commerce partnership marketing platform, high-scale distributed systems, social APIs
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Java or C# proficiency, OpenTelemetry implementation, Grafana ecosystem (Prometheus/Tempo/Loki), REST/Graph API integration, distributed tracing, capacity planning, root-cause analysis
Preferred skills
None stated
Technologies
OpenTelemetry, Grafana, Prometheus, Tempo, Loki, Jaeger, Honeycomb, Datadog, Java, C#, REST, OAuth
Responsibilities
Implement and maintain OpenTelemetry instrumentation across Java and C# services; Build integration tests with third-party social APIs and setup monitoring/alerting; Build and enhance Grafana dashboards tracking Golden Signals; Drive root-cause analysis for complex distributed system failures; Leverage tracing data to identify bottlenecks in cross-service communication; Debug issues across the entire stack from containerized code to cloud resources; Analyze application usage patterns to inform scaling decisions.
Seniority
Mid-to-Senior, hands-on IC