Staff Engineer, Site Reliability Engineering
Core
Design and operate scalable, fault-tolerant infrastructure and data platforms for vehicle telemetry and Software-Defined Vehicles, ensuring high reliability and observability.
Role type
Staff Site Reliability Engineer (SRE)
Builds
Cloud-native data ingestion pipelines, CI/CD delivery systems, and observability frameworks for automotive platforms.
Domain
Automotive / Software-Defined Vehicles / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
SRE/DevOps leadership, high-scale cloud-native system architecture, observability patterns (SLO/SLI), CI/CD pipeline design, Python/Go/Java programming, incident management, AI workflow automation, GitOps, infrastructure as code.
Preferred skills
Azure Databricks, Azure Event Hubs, AKS, Helm/Kustomize, Terraform, GitHub Actions/Argo CD, Prometheus/Grafana/Datadog/OpenTelemetry, LLM application development, Apache Flink/Kafka/Pulsar, vehicle telemetry systems.
Responsibilities
Lead design of scalable infrastructure for vehicle telemetry; define engineering patterns for service operability; manage production readiness and incident response; build and improve CI/CD pipelines; partner on SLO/SLI and monitoring strategies; automate operational workflows; mentor engineers and influence technical direction.
Seniority
Staff, hands-on technical leadership with team mentorship