Sr. SRE AI Engineer
Core
Building and operating reliable platforms, automation, and observability for AI-powered travel experiences, ensuring seamless service for business travelers and internal teams.
Role type
Senior Site Reliability Engineer (AI-focused)
Builds
Production platforms, automation tools, and observability systems for AI travel solutions
Domain
Travel technology and AI infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Java, Linux, Terraform, Grafana, Prometheus, SLOs, incident response, API reliability, rate limiting, quota management, authentication, latency optimization, root cause analysis, blameless postmortems
Preferred skills
AI provider integrations, AI agents, AI-assisted operational tools
Technologies
Terraform, CloudFormation, Grafana, Prometheus, New Relic, Datadog, Splunk
Responsibilities
Support AI-based application solutions and partner with development teams; Work with AI solutions, providers, and APIs to ensure reliability; Troubleshoot AI tools and provider issues; Operate reliable production platforms; Improve observability with dashboards and alerts; Apply AI to SRE workflows; Automate operational toil
Seniority
Senior, hands-on IC