Site Reliability Engineer, ARK Large Model Platform (Singapore)
Core
Develop and maintain the stability, observability, and cluster management of the Ark Large Model Platform on Volcano Engine to reduce IT costs for large model applications.
Role type
Site Reliability Engineer (Large Model Platform)
Builds
Large-scale model inference services and cloud-native resource scheduling systems
Domain
Cloud computing / Large-scale AI infrastructure
Deliverable
infrastructure
Required skills
Golang, Python, Java, cloud-native technologies, DevOps practices, observability systems, cluster management
Preferred skills
Infrastructure as Code (Terraform), stability systems for large-scale infrastructures, operating and maintaining large-scale systems
Technologies
VolcanoEngine, Terraform
Responsibilities
Manage and oversee the stability of control and data aspects of large-scale model systems; Develop and enhance observability systems for monitoring large model systems; Handle super large-scale cluster management and ensure efficient operation and maintenance.
Seniority
Mid-to-Senior, hands-on IC