Preferred Networks | Tokyo or Remote in Japan | Full-time | https://www.preferred.jp/en
Core
Build software infrastructure to serve LLMs using in-house inference accelerators and optimize inference engines for API services.
Role type
Senior IC systems engineer (LLM inference)
Builds
In-house LLM serving infrastructure and optimized inference engines for API services
Domain
AI infrastructure, LLMs, custom hardware acceleration
Deliverable
production ML models
Required skills
LLM serving, inference optimization, custom hardware acceleration, API service maintenance, open source contribution
Preferred skills
Experience with vLLM, MN-Core series, PLaMo series
Technologies
MN-Core, vLLM, PLaMo, Optuna, CuPy
Responsibilities
Build software infrastructure to serve LLMs using upcoming inference accelerator MN-Core L1000; Improve the inference engine powering API service; Maintain PLaMo implementations in open source projects
Seniority
Senior, hands-on IC
Sourced via hn_whoishiring · Listed on CareerPlan, which tracks 848,000+ jobs from 20+ sources.
