Staff Site Reliability Engineer
Core
Drive innovation and enhance service reliability by preventing repeatable issues, improving infrastructure scalability, and championing automation to eliminate manual activity.
Role type
Staff Site Reliability Engineer (AI-integration focus)
Builds
Reliable, scalable, and automated infrastructure systems for critical services
Domain
Cloud infrastructure, observability, and AI-driven workflow automation
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Python/Go/Java/Ruby coding, MySQL/PostgreSQL DBA, OpenTelemetry telemetry design, SLA management, system design at scale
Preferred skills
Observability at scale, DevOps automation (CI/CD), test automation, cloud engineering (Azure/AWS/GCP), Kubernetes orchestration, Ansible configuration management, Chaos engineering, incident response
Technologies
Python, Go, Java, Ruby, MySQL, PostgreSQL, OpenTelemetry, Gitlab CI-CD, Azure, AWS, GCP, Ansible, Kubernetes
Responsibilities
Proactively prevent repeatable issues using software development and systems engineering knowledge; Lead stakeholders to improve reliability and performance through system design; Champion automation culture to deliver scalable responses to system issues; Develop and maintain telemetry/monitoring solutions using OpenTelemetry; Collaborate with development teams to align new services with architectural standards
Seniority
Staff, hands-on IC with strategic influence