Senior Site Reliability Engineer
Core
Overseas model service operations, capacity management, and CI/CD automation for Tencent's Hunyuan AI systems.
Role type
Senior Site Reliability Engineer (Infrastructure & Cloud)
Builds
Stable, scalable overseas AI model services and automated operational tooling.
Domain
Artificial Intelligence / Cloud Infrastructure / Data Centers
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, network management, containerization, microservices architecture, public cloud operations (AWS/Azure), monitoring (Zabbix/Prometheus/Grafana), scripting (Python/Go/Shell), capacity planning, CI/CD automation.
Preferred skills
Experience with complex business system automation, industry trend exploration for intelligent O&M.
Technologies
Nginx, Redis, MySQL, Zabbix, Prometheus, Grafana, AWS, Azure, Python, Go, Shell, Kubernetes (implied by containerization), Docker (implied by containerization).
Responsibilities
Operate and maintain overseas model services for stability and efficiency; manage capacity and optimize resource costs; implement CI/CD pipelines and automated operational tools; design and improve service architectures; analyze system weaknesses using data-driven approaches; explore automation and intelligence trends in O&M.
Seniority
Senior, hands-on IC