大模型技术支持工程师-火山方舟(北京/杭州/成都)
Core
Monitor and troubleshoot large model training, inference, and platform services; handle on-call support for internal and external clients; manage large-scale batch inference operations.
Role type
Senior IC LLM Infrastructure Engineer
Builds
Large-scale batch inference tasks and platform services for LLMs
Domain
Artificial Intelligence / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes, cloud-native systems, public cloud platforms (AWS/GCP/Volcengine), GPU node health checks, Linux, networking, Python, Shell scripting, monitoring systems (Prometheus/Grafana)
Preferred skills
None stated
Technologies
Kubernetes, Prometheus, Grafana, Python, Shell, AWS, GCP, Volcengine
Responsibilities
Monitor and resolve alerts for model training and inference; act as on-call point of contact for technical support; manage instance scaling and traffic for batch inference; develop and maintain SOPs and automation tools
Seniority
Mid-level, hands-on IC