CareerPlanSign in

Token效能治理架构师 - TikTok研发

杭州💼 Full-time🗓 2026-09-28

Core

Design and govern the overall architecture for Token efficiency across the LLM inference lifecycle, covering collection, attribution, analysis, budgeting, optimization, evaluation, alerting, and operations.

Role type

Senior IC LLM inference cost and efficiency architect

Builds

Unified token metrics, inference optimization pipelines, cost attribution models, and governance SOPs for the TikTok AI platform

Domain

AI Infrastructure / LLM Inference / Cloud Cost Optimization

Deliverable

production ML models

Required skills

Distributed systems architecture, LLM inference mechanisms (Transformer, KV Cache, Batching, Quantization), Cost attribution modeling, Cross-team technical planning, Data-driven decision making, System performance tuning

Preferred skills

LLM inference engine development, Agent workload optimization, GPU/heterogeneous resource governance, Open source contributions

Responsibilities

Design end-to-end Token efficiency architecture; Establish unified metrics for token usage and cost; Optimize inference chains via prompt compression and caching; Govern Agent workloads to reduce token waste; Build evaluation guardrails for quality and safety; Drive cross-team technical planning and platform capability standardization

Sourced via bytedance · Listed on CareerPlan, which tracks 854,000+ jobs from 20+ sources.