Senior Software Engineer – Application Reliability , Hybrid
Core
Own the reliability of AI-powered applications and features from the user's perspective, focusing on feature uptime, automated diagnostics, and self-healing remediation.
Role type
Senior Software Engineer in Application Reliability (SRE)
Builds
LangGraph-based agents for automated diagnostics, Looker dashboards for observability, and evaluation harnesses for agent quality.
Domain
Generative AI, Cloud Infrastructure, Security
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python development, GCP (GKE, BigQuery, BigTable), SLI/SLO framework design, distributed tracing, SQL (complex queries, schema design), application-level debugging
Preferred skills
Agent evaluation harnesses, A2A protocols, streaming architectures, feature flags/canary deployments, GCP observability services, AIOps concepts, GenAI application development (LangGraph, prompt design), Looker dashboards
Technologies
LangGraph, Looker, BigQuery, BigTable, Python, GKE, Kubernetes
Responsibilities
Define and enforce feature-level SLIs, SLOs, and error budgets; Build and maintain application observability systems; Design and build LangGraph-based agents for automated issue identification and remediation; Develop agent evaluation harnesses; Write complex SQL for usage trend analysis and operational analytics; Analyze application usage trends to identify reliability risks; Partner with development teams to embed reliability practices; Lead application-level incident response and postmortems; Build Python-based tooling to reduce MTTD and MTTR.
Seniority
Senior, hands-on IC