Senior Site Reliability Engineer / VMWare Expert (GDL - CDMX)
Core
Plan, monitor, and optimize capacity across hybrid infrastructure (VMware and public cloud) while ensuring high availability, scalability, and performance of enterprise platforms.
Role type
Senior Site Reliability Engineer (VMware & Hybrid Cloud)
Builds
Hybrid cloud infrastructure, containerized workloads, and automated operational pipelines
Domain
Cloud Infrastructure / Virtualization / FinOps
Deliverable
production ML models | infrastructure
Required skills
VMware vSphere administration, Hybrid cloud management, Kubernetes orchestration, Infrastructure as Code, CI/CD pipeline management, Linux performance tuning, Observability tooling, Scripting for automation, Capacity planning, Root cause analysis
Preferred skills
Cloud FinOps, Disaster Recovery planning, Predictive analytics, Serverless architectures, AI tooling for operations
Technologies
VMware vSphere, AWS, Azure, GCP, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI, GitHub Actions, Azure DevOps, Dynatrace, Datadog, Prometheus, Grafana, New Relic, Zabbix, Python, Bash, PowerShell, PowerCLI
Responsibilities
Plan and optimize capacity across hybrid infrastructure; Ensure high availability and scalability of enterprise platforms; Implement and maintain observability solutions; Analyze capacity trends and provide forecasting; Collaborate with Architecture, DevOps, and Cloud teams; Identify performance bottlenecks and drive optimization strategies; Develop automation for infrastructure provisioning and operations; Support Kubernetes and containerized workloads; Lead root cause analysis and troubleshoot complex issues; Produce capacity planning and executive reporting
Seniority
Senior, hands-on IC