Software Engineer, Site Reliability
Core
Build and operate systems to improve the reliability, resiliency, and observability of Upstart's production environment, integrating AI into the engineering lifecycle.
Role type
Software Engineer, Site Reliability Engineering (SRE)
Builds
Shared observability capabilities, incident response systems, operational automation, and resiliency tools for production systems.
Domain
Fintech / AI Lending / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, JavaScript, TypeScript, distributed systems, cloud infrastructure, observability, incident response, AI-assisted development, automation
Preferred skills
Kubernetes, AWS, infrastructure as code, Datadog, Sumo Logic, CloudWatch, service level objectives, capacity planning, disaster recovery, AI-enabled operational workflows
Technologies
Kubernetes, AWS, Datadog, Sumo Logic, CloudWatch
Responsibilities
Build and improve tooling for system reliability and observability; Develop shared observability capabilities for metrics, logs, and traces; Improve incident response and operational readiness; Build resiliency capabilities to identify failure modes and recover from failures; Automate recurring operational toil; Use AI to enhance incident investigation and engineering efficiency
Seniority
Mid-level, hands-on IC