大模型应用架构工程师 - 火山方舟
Core
Design and optimize LLM inference architecture for To B and To C scenarios, ensuring low latency and cost efficiency while managing global infrastructure and heterogeneous hardware.
Role type
Senior LLM Application Architect (Inference & Performance)
Builds
High-throughput, stable LLM inference services and AI application platforms
Domain
Large Language Models, Distributed Systems, Cloud Infrastructure
Deliverable
production ML models
Required skills
C++, Python, Rust, Distributed Systems, Heterogeneous Inference, System Stability, Cost Optimization, Framework Design
Preferred skills
Global Architecture Design, Automated Engineering, Algorithmic Optimization
Technologies
LLMs, Heterogeneous Hardware, Cloud Infrastructure
Responsibilities
Implement LLM inference solutions for diverse business scenarios; Optimize inference performance, stability, and throughput; Manage global architecture and heterogeneous hardware adaptation.