Software Engineering, MTS (SRE & Devops)
Core
Build and manage a multi-substrate Kubernetes and microservices platform powering Core CRM and growing applications, ensuring high availability and reliability for tens of millions of users.
Role type
Senior Site Reliability Engineer (SRE) / DevOps
Builds
A distributed systems engineering platform for Salesforce Core CRM and applications
Domain
Cloud Infrastructure / SaaS / Kubernetes
Deliverable
infrastructure
Required skills
Linux systems administration, Kubernetes, Python, Go, Terraform, automation scripting, distributed systems architecture, troubleshooting complex production issues, monitoring and metrics implementation, self-healing mechanisms design, networking protocols (TCP/IP, Load Balancers)
Preferred skills
Experience with service mesh, experience with Spinnaker, experience with Puppet/Chef/Ansible, experience with Nagios/Grafana/Zabbix, experience with AWS
Technologies
Kubernetes, Docker, Spinnaker, Puppet, Chef, Ansible, Nagios, Grafana, Zabbix, AWS, Terraform, Jenkins, Python, GoLang
Responsibilities
Ensure high availability of large fleet of clusters running Kubernetes and related technologies, troubleshoot real production issues, contribute code to drive platform improvements, drive automation efforts to eliminate manual work, implement monitoring and metrics for platform visibility, implement self-healing mechanisms to proactively fix issues, evaluate new technologies to solve problems
Seniority
Mid-Senior, hands-on IC