Cloud Engineer II
Core
Design, build, and maintain a scalable Service Mesh and Kubernetes architecture on AWS to ensure high availability, performance, and reliability for the global TraceLink platform.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Scalable cloud infrastructure, Service Mesh, and Kubernetes clusters for multi-enterprise supply chain operations
Domain
Cloud Infrastructure, Kubernetes, Service Mesh, Multi-cloud (AWS/Azure)
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform, Helm, AWS services, Python, Shell scripting, container diagnostics, multi-cloud management, cost optimization
Preferred skills
Observability, Istio, Agentic AI capabilities
Technologies
AWS, Azure, Kubernetes, Terraform, Helm, Python, Shell, Istio
Responsibilities
Design and build new tools and technologies to eliminate bugs and increase performance; monitor systems to proactively detect and address issues; gather and analyze metrics to assist in performance tuning and fault finding; optimize AWS resources usage across multiple environments; collaborate with engineering stakeholders to improve service availability and reliability
Seniority
Senior, hands-on IC
