Software Developer(SRE) - CloudVision as a Service (CVaaS)
Core
Operate and scale the global CloudVision service fleet on Kubernetes, focusing on reliability, observability, and automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Global CloudVision service fleet, CI/CD pipelines, and automated incident response systems
Domain
Cloud networking, Kubernetes, distributed systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Linux administration, Go, Python, bash shell scripting, infrastructure-as-code, distributed database management, real-time stream processing, observability stack management
Preferred skills
PostgreSQL, Docker, virtualization, Artifactory, GitLab, Spinnaker, Terraform
Technologies
Kubernetes, GKE, Spinnaker, HBase, Hadoop, ElasticSearch, ClickHouse, Kafka, TensorFlow, Prometheus, Grafana, Loki
Responsibilities
Build, deploy, and operate critical production systems with focus on scalability and reliability; Monitor, support, and enhance product deployment experience; Build automation to remove toil; Proactively monitor, respond to, and enhance alerts; Create and maintain incident response runbooks; Triage platform/infrastructural issues and engage with 3rd party vendor support; Plan and communicate maintenance windows; Work with product development teams to identify and resolve infrastructural bottlenecks; Survey and adopt best practices for secure, scalable, and fault-tolerant systems
Seniority
Senior, hands-on IC
