Senior Software Engineer - SRE
Core
Design, deliver, and maintain a highly available and performant application platform while building observability tools and automating operational processes.
Role type
Senior Software Engineer - Site Reliability Engineering (SRE)
Builds
Mission-critical application platforms, observability tools, and AI-assisted incident response systems
Domain
Cloud infrastructure, distributed systems, and AI/ML for reliability engineering
Deliverable
production ML models | infrastructure
Required skills
Java, Spring, cloud infrastructure (Azure/AWS/GCP), JVM tuning, memory analysis, Kubernetes, distributed systems, microservices, CI/CD, Terraform, Helm, Jenkins, GitLab, Python, Bash, SQL/NoSQL, Datadog, Prometheus, Grafana, LLM integration, vector databases, RAG architectures, prompt engineering
Preferred skills
chaos engineering, incident management platforms (PagerDuty), product-facing high-traffic services
Technologies
Java, Spring, Azure, AWS, GCP, Kubernetes, EKS, AKS, GKE, Datadog, Prometheus, Grafana, Terraform, Helm, Jenkins, GitLab, Python, Bash, PagerDuty
Responsibilities
Design and instrument SLIs/SLOs for latency, error rates, and availability; manage error budgets; improve alert quality by reducing noise; build scripts for operational automation and incident response; develop or integrate LLM-based tools to reduce MTTR; apply machine learning for anomaly detection and capacity prediction; embed with product teams to review architectures and catch reliability risks
Seniority
Senior, hands-on IC