Sr. Software Engineer - Distributed Systems
Core
Design, build, and improve large-scale distributed observability services (Monitoring, Logging, Alerting, Tracing) for AI-driven platforms, handling hundreds of terabytes of data and billions of daily messages.
Role type
Senior Software Engineer (Distributed Systems & Observability)
Builds
Real-time and batch data processing pipelines, distributed tracing systems, metrics ingestion, and AI-driven observability tools (anomaly detection, automated root-cause analysis).
Domain
Cloud-native infrastructure, Distributed Systems, AI Observability
Deliverable
production ML models | product features | infrastructure
Required skills
Python, Go, or Java; Linux; Distributed tracing; Time-Series DBs; Containerization (Docker, Kubernetes); Cloud-native tools (Prometheus, Service Mesh); LLM orchestration frameworks (LangChain, LlamaIndex, Semantic Kernel); AI system debugging.
Preferred skills
MS Degree; Experience with AWS/GCP native observability tooling; Experience with infrastructure automation (Terraform, Ansible, Chef).
Technologies
Kubernetes, Docker, OpenStack, Prometheus, AWS CloudWatch, GCP Cloud Operations, LangChain, LlamaIndex, Semantic Kernel, Python, Go, Java.
Responsibilities
Design and develop core software modules for real-time/batch data processing; Build data capture services across multiple infrastructure types; Own distributed tracing end-to-end; Instrument AI-powered workflows for performance and cost; Partner with product teams to define SLOs/SLIs; Evaluate and implement new open-source/cloud-native tools; Participate in on-call rotation.
Seniority
Senior, hands-on IC