硬件加速推理AI Infra工程师 - 芯片研发
Core
Adapt large model business models to self-developed chips, perform hardware-software co-optimization, and build distributed inference systems to maximize hardware utilization and deployment throughput.
Role type
Senior IC hardware acceleration AI infrastructure engineer (chip R&D)
Builds
Custom hardware optimizations for Byte's proprietary business scenarios, distributed inference systems, and AI model deployment frameworks.
Domain
Semiconductor industry + AI infrastructure
Deliverable
production ML models
Required skills
AI accelerator architecture and parallel computing, C/C++ and Python, ONNX/TensorFlow/PyTorch, model quantization and sparsity, distributed system design
Preferred skills
Compiler development (MLIR/TVM), GPU/AI chip architecture and operator optimization, LLM/multimodal model expertise, AI server cluster architecture, vLLM/SGLang framework development
Responsibilities
Evaluate adaptability and performance of large models on self-developed chips, optimize AI model compilation frameworks for hardware utilization, tune distributed inference systems for high throughput, implement model quantization and distillation solutions
Seniority
Senior, hands-on IC
