Staff Software Engineer, Metrics and Logging
Core
Designing and scaling the next-generation logging platform to process petabytes of logs daily, enabling deep system insights and efficient troubleshooting across Databricks services.
Role type
Staff Software Engineer (Infrastructure/Logging)
Builds
Scalable, low-latency log delivery pipelines and observability tools for a global data and AI infrastructure platform.
Domain
Cloud infrastructure, distributed systems, observability, data engineering.
Deliverable
production ML models | infrastructure
Required skills
Scala, Rust, Go, Python, Java, C++, large-scale distributed systems, log collection, health monitoring, observability tools, complex project leadership, cross-team collaboration.
Preferred skills
Structured logging best practices, cost-efficiency optimization, retention and indexing strategies.
Technologies
Apache Spark, Delta Lake, MLflow, petabyte-scale data processing.
Responsibilities
Design and scale the next-generation logging platform; develop and optimize log delivery pipelines for high-throughput ingestion; enhance log accessibility and usability tools; define best practices for structured logging; improve reliability and cost-efficiency of log retention and querying; mentor engineers and foster technical excellence.
Seniority
Staff, hands-on IC with strategic impact and mentorship.