边缘基础设施运维工程师
Core
Manage the full lifecycle of physical edge infrastructure (GPU/CPU servers, storage) and operate Kubernetes-based container platforms for external customer compute delivery.
Role type
Senior Infrastructure Operations Engineer (Edge Computing & GPU)
Builds
Edge data centers, GPU compute clusters, and containerized service delivery pipelines for external clients.
Domain
Edge Computing, GPU Infrastructure, Cloud Native
Deliverable
production ML models | infrastructure
Required skills
GPU server lifecycle management, Kubernetes cluster operations, network configuration (VLAN/VXLAN/BGP), Linux system tuning, Python/Go/Shell scripting, Prometheus/Grafana monitoring, hardware fault diagnosis
Preferred skills
Edge computing scenarios, Terraform/Ansible IaC, GPU resource pooling (MIG/vGPU), toB platform operations, domestic GPU ecosystem (Ascend/Huawei)
Technologies
Kubernetes, Docker, containerd, NVIDIA A100/H100/L40S, GPU Operator, Prometheus, Grafana, Loki, ELK, CSI, Service Mesh, Terraform, Ansible
Responsibilities
Manage physical edge server hardware (GPU/CPU/storage) including deployment, firmware upgrades, and fault replacement; Configure and monitor edge network devices (switches/routers/firewalls) to ensure low latency; Operate and scale Kubernetes edge clusters with GPU resource scheduling; Build and maintain customer-facing compute delivery pipelines (Helm, quotas, isolation); Develop automation scripts and tools for infrastructure as code; Monitor SLO/SLA metrics and respond to production incidents.
Seniority
Senior, hands-on IC