Senior / Staff Software Engineer (Observability / SRE)
Core
Design and lead the architecture of monitoring and observability stacks for cloud and on-prem environments, optimizing end-to-end performance across hardware, firmware, and distributed services.
Role type
Senior/Staff Software Engineer (Observability / SRE)
Builds
Production observability tooling, CI/CD-based performance regression detection pipelines, and telemetry/alerting systems.
Domain
Autonomous transportation, Physical AI, cloud infrastructure, and distributed systems.
Deliverable
production ML models | infrastructure
Required skills
Python, Rust, C/C++, Linux internals, perf tooling, Kubernetes, microservices, distributed systems, system design, end-to-end project ownership.
Preferred skills
OpenTelemetry, Grafana OSS, GPU/xPU performance tuning, Argo Workflows, PyTorch, distributed tracing, Prometheus.
Technologies
Go, Python, Java, Kubernetes, Docker, perf, eBPF, flamegraphs, OpenTelemetry, Grafana, Prometheus, Argo Workflows, PyTorch.
Responsibilities
Design and lead observability stack architecture; develop workloads and benchmarks for compute/storage/network/ML; analyze and optimize performance using profiling tools; build automation and observability tooling; support client teams' observability requirements; influence system architecture decisions; drive execution, write design docs, and mentor ICs.
Seniority
Senior/Staff, hands-on IC with leadership and mentorship responsibilities.