Member of Technical Staff - Applied ML, Japanese Multimodal
Core
Deploy general-purpose AI systems (LLMs, multimodal models) for enterprise customers in Japan, optimizing for latency, memory, and reliability across edge and data center environments.
Role type
Senior Applied ML Engineer (Deployment & Optimization)
Builds
Production-ready AI solutions, inference pipelines, evaluation systems, and surrounding software for customer integration.
Domain
Artificial Intelligence / Machine Learning / Enterprise Software
Deliverable
production ML models
Required skills
Production ML system deployment, model inference optimization, post-training (SFT, PEFT), model evaluation & error analysis, open-source ML ecosystem, technical customer collaboration, English proficiency
Preferred skills
Japanese proficiency, LLM post-training methods, inference frameworks (vLLM, SGLang, llama.cpp, ONNX Runtime, MLX), quantization, edge/mobile/embedded deployment, multimodal systems
Technologies
vLLM, SGLang, llama.cpp, ONNX Runtime, MLX, LLMs, multimodal models
Responsibilities
Own end-to-end applied ML projects from discovery to production deployment; Integrate and optimize model inference for latency, throughput, and memory constraints; Build data pipelines and serving components; Fine-tune or post-train models; Design evaluations and conduct error analysis; Collaborate with customer engineering teams on design and rollout.
Seniority
Senior, hands-on IC