SRE Engineer
Core
Ensuring reliability, performance, and automation excellence across the Ocean & Enablement Platform container technology ecosystem.
Role type
Senior IC Site Reliability Engineer (SRE)
Builds
Reliable, cost-efficient, and automated multi-cloud container infrastructure for global logistics operations.
Domain
Logistics / Cloud Infrastructure / AIOps
Deliverable
production ML models | infrastructure
Required skills
Python scripting, AWS/Azure/GCP, Docker, Kubernetes, Terraform, Ansible, Prometheus, Grafana, ELK stack, incident response, root cause analysis, SLI/SLO definition
Preferred skills
AIOps, AI-driven automation, LLM platforms, cloud cost optimization, CI/CD pipelines, relational databases
Technologies
Python, Ansible, shell scripting, AWS, Azure, GCP, Docker, Kubernetes, Terraform, Prometheus, Grafana, ELK, Datadog, Azure AI Foundry, Azure OpenAI
Responsibilities
Support and improve reliability, availability, and performance across O&E applications and shared services; Participate in on-call rotations, handle incidents, perform RCA, and implement corrective actions; Develop and maintain automation tools and scripts using Python, Ansible, or shell scripting to reduce manual toil; Deploy, monitor, and manage workloads on AWS / Azure / GCP, ensuring cost-efficient and reliable operations; Configure and enhance observability (Prometheus, Grafana, ELK) for proactive detection and fast recovery; Support SRE principles by defining SLIs, SLOs, and error budgets for O&E services; Explore AIOps and AI-driven automation to reduce alert noise, accelerate triage, and enable intelligent remediation; Collaborate with product, infra, and observability teams to improve reliability and incident response; Maintain runbooks, SOPs, and automation playbooks for operational readiness.
Seniority
Mid-level, hands-on IC