Staff Software Engineer, Observability
Required skills
BS (or higher) in Computer Science, or a related field. 7+ years of production-level experience in one of: Go, Python, Java, Scala, Rust, C++, or similar languages. Experience in software development, in large-scale distributed systems. Experience driving large projects involving multiple teams. Experience with cloud technologies, e.g. AWS, Azure, GCP, Docker, or Kubernetes. Familiarity with observability infrastructure, monitoring patterns, and reliability practices.
Responsibilities
Build the next generation of observability platforms that support billions of active time series and process petabytes of logs daily. Manage infrastructure across nearly a hundred cloud regions, enabling all Databricks engineers and customers to monitor the reliability of our product. Develop advanced workflows that accelerate incident diagnosis for Bricksters, allowing engineers to quickly derive insights from logs and metrics. You will leverage powerful capabilities of Databricks’ own data intelligence platform to push the boundaries of troubleshooting practices in the industry. Uplevel monitoring and reliability practices across Databricks engineering, developing opinionated tools that set common standards for managing structured logs, metrics, alerts, dashboards, and oncall rotations. Mentor and uplevel engineers, fostering a culture of technical excellence within the team and broader observability community.
Seniority
...
Domain
Observability, Data and AI infrastructure platform
Full job description
- Build the next generation of observability platforms that support billions of active time series and process petabytes of logs daily.
- Manage infrastructure across nearly a hundred cloud regions, enabling all Databricks engineers and customers to monitor the reliability of our product.
- Develop advanced workflows that accelerate incident diagnosis for Bricksters, allowing engineers to quickly derive insights from logs and metrics. You will leverage powerful capabilities of Databricks’ own data intelligence platform to push the boundaries of troubleshooting practices in the industry.
- Uplevel monitoring and reliability practices across Databricks engineering, developing opinionated tools that set common standards for managing structured logs, metrics, alerts, dashboards, and oncall rotations.
- Mentor and uplevel engineers, fostering a culture of technical excellence within the team and broader observability community.
Requirements
- BS (or higher) in Computer Science, or a related field.
- 7+ years of production-level experience in one of: Go, Python, Java, Scala, Rust, C++, or similar languages.
- Experience in software development, in large-scale distributed systems.
- Experience driving large projects involving multiple teams.
- Experience with cloud technologies, e.g. AWS, Azure, GCP, Docker, or Kubernetes.
- Familiarity with observability infrastructure, monitoring patterns, and reliability practices.
Nice to Have
- N/A
Benefits
- Competitive salary and total compensation package including annual performance bonus, equity, and benefits.
- Opportunity to work on large-scale, high-impact projects.
- Collaborative and innovative work environment.
- Global offices and remote work options.