Staff Software Engineer in Hardware Infrastructure Observability
Core
Building low-level monitoring, metrics, and maintenance automation for large-scale server fleets and data center engineering systems to ensure reliability and enable safe fleet-wide changes.
Role type
Staff Software Engineer in Hardware Infrastructure Observability
Builds
Monitoring services, agents, metrics pipelines, alerting systems, and maintenance automation workflows for data center infrastructure.
Domain
Cloud infrastructure, data center operations, hardware observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python, Golang, Linux fundamentals, incident investigation, system design, stakeholder management, architectural thinking, mentoring
Preferred skills
Server architecture knowledge, Prometheus-compatible stacks (e.g., VictoriaMetrics), computer networking, high-load distributed systems
Technologies
Python, Golang, Linux, Prometheus, VictoriaMetrics
Responsibilities
Design and develop services and agents for deep visibility into server fleets; evolve metrics/aggregation/alerting pipelines; build maintenance workflows and automation; investigate incidents hands-on; collaborate with hardware and DC operations teams.
Seniority
Staff, hands-on IC with mentorship and architectural leadership