TIONE-高级后台研发工程师-上海/杭州/深圳/北京
Core
Building a platform for large model inference services on Tencent Cloud, focusing on model deployment, service orchestration, and performance optimization.
Role type
Senior backend engineer (LLM inference platform)
Builds
Scalable, high-availability inference services for large language models
Domain
Cloud computing, AI infrastructure, distributed systems
Deliverable
production ML models
Required skills
Go, Python, Linux, Kubernetes, Docker, high-concurrency system design, gRPC, HTTP, SSE, GPU cluster management, SLO management, capacity planning
Preferred skills
vLLM, SGLang, PagedAttention, continuous batching, TP/PP parallelism, PD separation, quantization, long context handling
Responsibilities
Design and implement platform features for model access, versioning, and deployment; Optimize inference engines for throughput and latency; Ensure system stability through observability and fault analysis; Adapt infrastructure for heterogeneous compute resources.