CareerPlanSign in

大模型应用架构工程师 - 火山方舟

北京💼 Full-time🗓 2026-09-28

Core

Design and optimize LLM inference architecture for To B and To C scenarios, ensuring low latency and cost efficiency while managing global infrastructure and heterogeneous hardware.

Role type

Senior LLM Application Architect (Inference & Performance)

Builds

High-throughput, stable LLM inference services and AI application platforms

Domain

Large Language Models, Distributed Systems, Cloud Infrastructure

Deliverable

production ML models

Required skills

C++, Python, Rust, Distributed Systems, Heterogeneous Inference, System Stability, Cost Optimization, Framework Design

Preferred skills

Global Architecture Design, Automated Engineering, Algorithmic Optimization

Technologies

LLMs, Heterogeneous Hardware, Cloud Infrastructure

Responsibilities

Implement LLM inference solutions for diverse business scenarios; Optimize inference performance, stability, and throughput; Manage global architecture and heterogeneous hardware adaptation.

Sourced via bytedance · Listed on CareerPlan, which tracks 852,000+ jobs from 20+ sources.