Software Engineering Systems Engineer
Core
Ensure availability and performance for Tableau Online cloud service by leading incident response and driving AI-driven operational solutions.
Role type
Senior Site Reliability Operations Engineer (SRE)
Builds
Tableau Online cloud products and services
Domain
SaaS / Cloud Infrastructure / Data Analytics
Deliverable
production ML models | infrastructure
Required skills
Incident command and leadership, Kubernetes, Terraform, Splunk, Grafana, distributed tracing, Python, Java, AWS, AI tools for automation, Agile methodologies
Preferred skills
Open source contributions, coaching team members, process evaluation
Technologies
Kubernetes, Terraform, Spinnaker, Splunk, Grafana, AWS, Python, Java, Claude, Resolve.ai
Responsibilities
Lead incident responders to mitigate customer incidents, adopt best practices to improve availability and resiliency, utilize automation and AI to reduce operational toil, provide technical leadership at component scope, evaluate and implement process changes for efficiency
Seniority
Senior, hands-on IC with leadership