Site Reliability Engineer III
Core
Building automation to improve system reliability and resilience for large-scale, distributed, fault-tolerant cloud platforms.
Role type
Senior IC Site Reliability Engineer
Builds
Cloud-based transformations and large-scale distributed software applications
Domain
Cloud infrastructure (GCP) and distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, containers, elastic scalability, SRE principles, API/microservice architecture, infrastructure/network/database/OS troubleshooting, production deployment, CI/CD, performance diagnostics, capacity planning, performance tuning, web technologies (HTTP, proxy, Java)
Preferred skills
Dynatrace, Azure DevOps, Prometheus, Terraform, Grafana
Technologies
Google Cloud Platform (GCP), Kubernetes, Terraform, Prometheus, Grafana, Dynatrace, Azure DevOps, Java
Responsibilities
Gathers and analyzes metrics from monitoring platforms for performance tuning and fault tolerance; Partners with development teams to improve services through testing and release procedures; Participates in system design, platform management and capacity planning; Balances feature development speed and reliability with service-level objectives; Works closely with the incident response team to restore service; Investigates, blocks and rate-limits unwanted traffic; Utilizes monitoring systems and dashboards for proactive changes and alerting; Establishes continuous process improvement cycles.
Seniority
Senior, hands-on IC