Inference Software Engineer
Core
Designing and implementing high-performance software stacks for AI inference on custom ASICs, enabling real-time video generation and deep reasoning agents.
Role type
Inference Software Engineer (AI Hardware Stack)
Builds
High-throughput inference serving front-end and host software stack for model-specific ASICs
Domain
AI Infrastructure / Custom Hardware Acceleration
Deliverable
production ML models
Required skills
C++, Python, transformer model architectures, inference serving stacks, distributed inference environments
Preferred skills
Rust, GPU kernels, CUDA compilation stack, distributed systems, networking, parallel programming
Technologies
Rust, C++, Python, vLLM, SGLang, CUDA
Responsibilities
Contribute to architecture and design of the Sohu host software stack; Implement high-performance, modular code across the complete Etched software stack; Interface with firmware and drivers teams delivering highest-performance HW/SW stack; Work with AI model researchers and product-facing teams building out the Etched serving front-end
Seniority
Mid-Senior, hands-on IC