Token效能治理架构师 - TikTok研发
Core
Design and govern the overall architecture for Token efficiency across the LLM inference lifecycle, covering collection, attribution, analysis, budgeting, optimization, evaluation, alerting, and operations.
Role type
Senior IC LLM inference cost and efficiency architect
Builds
Unified token metrics, inference optimization pipelines, cost attribution models, and governance SOPs for the TikTok AI platform
Domain
AI Infrastructure / LLM Inference / Cloud Cost Optimization
Deliverable
production ML models
Required skills
Distributed systems architecture, LLM inference mechanisms (Transformer, KV Cache, Batching, Quantization), Cost attribution modeling, Cross-team technical planning, Data-driven decision making, System performance tuning
Preferred skills
LLM inference engine development, Agent workload optimization, GPU/heterogeneous resource governance, Open source contributions
Responsibilities
Design end-to-end Token efficiency architecture; Establish unified metrics for token usage and cost; Optimize inference chains via prompt compression and caching; Govern Agent workloads to reduce token waste; Build evaluation guardrails for quality and safety; Drive cross-team technical planning and platform capability standardization