Senior Site Reliability Engineer
Core
Designing, implementing, and operating cloud infrastructure and automation tooling to ensure high availability and reliability for Navan's travel and expense services used by thousands of daily travelers.
Role type
Senior Site Reliability Engineer (IC)
Builds
Autonomous, monitored, fault-tolerant cloud infrastructure and AI/ML microservices for enterprise travel services
Domain
Travel technology / Cloud Infrastructure / AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
Java (JVM profiling), Python, Bash, Go, Terraform, AWS, CI/CD, Microservices architecture, SLO/SLI frameworks, Observability (NewRelic, DataDog, Kibana), AI/LLM platform management, Infrastructure as Code
Preferred skills
Node.js, CloudFormation, RAG pipelines, Agentic systems, AWS Fargate/ECS, Gemini, Braintrust, Predictive insights via telemetry
Technologies
Java, Python, Bash, Go, Node.js, Terraform, CloudFormation, AWS, Jenkins, Maven, NewRelic, DataDog, SignalFX, Kibana, AWS Fargate, AWS ECS, Gemini, Braintrust
Responsibilities
Design and develop tooling, automation, and infrastructure services; Identify reliability anti-patterns and solve them systemically; Automate toil in processes; Define and drive adoption of system reliability standards (SLO/SLI, blameless post-mortems); Deploy and monitor production-grade AI/ML microservices; Integrate AI-specific telemetry for root-cause analysis
Seniority
Senior, hands-on IC with mentorship responsibilities