大模型基础设施工程师(大模型资源管理/数据管理处理方向)-TikTok Shop
Core
Design and implement multi-tenant compute resource management and scheduling systems for e-commerce scenarios, covering GPU/CPU/memory/network, and build end-to-end data platforms for large model training, fine-tuning, and evaluation.
Role type
Senior Infrastructure Engineer (Large Model Resource & Data Management)
Builds
Scalable compute clusters for training/inference, FinOps cost optimization systems, and high-throughput data pipelines for SFT/RLHF.
Domain
E-commerce, Large Language Models (LLM), Cloud Infrastructure
Deliverable
infrastructure
Required skills
Go/Java/Python, Kubernetes, GPU scheduling, Distributed systems, Object storage, Stream processing (Kafka/Flink), Observability (Prometheus/Grafana), FinOps
Preferred skills
NCCL/CUDA, Triton/Ray, RLHF platform engineering, Multi-cloud practices, RDMA/NVLink optimization
Technologies
Kubernetes, Volcano, Kueue, Iceberg, Delta Lake, Kafka, Flink, Spark, Prometheus, Grafana, ELK, IAM, KMS
Responsibilities
Design multi-tenant resource scheduling and quota isolation; Optimize cluster scheduling strategies for SLA stability; Build FinOps capabilities for cost attribution and optimization; Implement auto-scaling and disaster recovery; Construct data platforms for data lineage and compliance; Build batch/stream ETL pipelines with quality monitoring.
Seniority
Senior, hands-on IC