CareerPlanSign in

Member of Technical Staff - Applied ML, Japanese Multimodal

Tokyo🌐 Remote💼 Full-time🗓 2026-07-29 → 2026-09-27

Core

Deploy general-purpose AI systems (LLMs, multimodal models) for enterprise customers in Japan, optimizing for latency, memory, and reliability across edge and data center environments.

Role type

Senior Applied ML Engineer (Deployment & Optimization)

Builds

Production-ready AI solutions, inference pipelines, evaluation systems, and surrounding software for customer integration.

Domain

Artificial Intelligence / Machine Learning / Enterprise Software

Deliverable

production ML models

Required skills

Production ML system deployment, model inference optimization, post-training (SFT, PEFT), model evaluation & error analysis, open-source ML ecosystem, technical customer collaboration, English proficiency

Preferred skills

Japanese proficiency, LLM post-training methods, inference frameworks (vLLM, SGLang, llama.cpp, ONNX Runtime, MLX), quantization, edge/mobile/embedded deployment, multimodal systems

Technologies

vLLM, SGLang, llama.cpp, ONNX Runtime, MLX, LLMs, multimodal models

Responsibilities

Own end-to-end applied ML projects from discovery to production deployment; Integrate and optimize model inference for latency, throughput, and memory constraints; Build data pipelines and serving components; Fine-tune or post-train models; Design evaluations and conduct error analysis; Collaborate with customer engineering teams on design and rollout.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 814,000+ jobs from 20+ sources.