CareerPlanSign in

Senior Site Reliability Engineer

Singapore-CapitaSky💼 Full-time🗓 2025-12-24 → 2026-09-26

Core

Overseas model service operations, capacity management, and CI/CD automation for Tencent's Hunyuan AI systems.

Role type

Senior Site Reliability Engineer (Infrastructure & Cloud)

Builds

Stable, scalable overseas AI model services and automated operational tooling.

Domain

Artificial Intelligence / Cloud Infrastructure / Data Centers

Deliverable

production ML models | infrastructure

Required skills

Linux system administration, network management, containerization, microservices architecture, public cloud operations (AWS/Azure), monitoring (Zabbix/Prometheus/Grafana), scripting (Python/Go/Shell), capacity planning, CI/CD automation.

Preferred skills

Experience with complex business system automation, industry trend exploration for intelligent O&M.

Technologies

Nginx, Redis, MySQL, Zabbix, Prometheus, Grafana, AWS, Azure, Python, Go, Shell, Kubernetes (implied by containerization), Docker (implied by containerization).

Responsibilities

Operate and maintain overseas model services for stability and efficiency; manage capacity and optimize resource costs; implement CI/CD pipelines and automated operational tools; design and improve service architectures; analyze system weaknesses using data-driven approaches; explore automation and intelligence trends in O&M.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.