Software Engineer, Reliability Platforms
Core
Design, build, and operate services and infrastructure for service health, change orchestration, and incident management to enable 4,000+ internal customers.
Role type
Software Engineer, Reliability Platforms
Builds
SLO frameworks, analytics tools, AI Agent enablement, self-service provisioning orchestration, incident management tools, and runtime configuration tooling.
Domain
Food delivery / Cloud Infrastructure / Observability
Deliverable
production ML models | infrastructure
Required skills
Software development, systems integration, SRE mindset, AI/LLM integration, orchestration workflows, backend services design, telemetry analysis, automation engineering
Preferred skills
Creative pioneering mindset, prototype development, frontier thinking, risk balancing
Technologies
Kafka, Databases, MCP (Model Context Protocol), AI Agents, Kubernetes (implied by pods), feature flag systems
Responsibilities
Design and deliver SLO quality frameworks for tens of thousands of endpoints; Build automated alert routing and escalation management tools; Develop backend services for reliability platform data and tools; Create orchestration tools for self-service provisioning of critical infrastructure; Implement per-pod realtime configuration key-value tooling for runtime feature flags; Contribute to AI Agentic tooling for troubleshooting and Q&A.
Seniority
Individual Contributor (IC), hands-on engineering