CareerPlanSign in

大模型基础设施工程师(大模型资源管理/数据管理处理方向)-TikTok Shop

上海💼 Full-time🗓 2026-09-28

Core

Design and implement multi-tenant compute resource management and scheduling systems for e-commerce scenarios, covering GPU/CPU/memory/network, and build end-to-end data platforms for large model training, fine-tuning, and evaluation.

Role type

Senior Infrastructure Engineer (Large Model Resource & Data Management)

Builds

Scalable compute clusters for training/inference, FinOps cost optimization systems, and high-throughput data pipelines for SFT/RLHF.

Domain

E-commerce, Large Language Models (LLM), Cloud Infrastructure

Deliverable

infrastructure

Required skills

Go/Java/Python, Kubernetes, GPU scheduling, Distributed systems, Object storage, Stream processing (Kafka/Flink), Observability (Prometheus/Grafana), FinOps

Preferred skills

NCCL/CUDA, Triton/Ray, RLHF platform engineering, Multi-cloud practices, RDMA/NVLink optimization

Technologies

Kubernetes, Volcano, Kueue, Iceberg, Delta Lake, Kafka, Flink, Spark, Prometheus, Grafana, ELK, IAM, KMS

Responsibilities

Design multi-tenant resource scheduling and quota isolation; Optimize cluster scheduling strategies for SLA stability; Build FinOps capabilities for cost attribution and optimization; Implement auto-scaling and disaster recovery; Construct data platforms for data lineage and compliance; Build batch/stream ETL pipelines with quality monitoring.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.