ML Solution Architect (Early Talent)
Core
Building and testing LLM-based solutions and applications using Nebius Token Factory's serverless inference and fine-tuning platform for open-source LLMs.
Role type
Early career ML Solutions Architect (hands-on IC)
Builds
LLM-based solutions and applications (multimodal: text, vision, audio) on the Token Factory platform
Domain
Generative AI / Cloud Infrastructure / LLM Inference
Deliverable
production ML models
Required skills
Python programming, Generative AI development, ML frameworks (PyTorch, Transformers), Prompt engineering, Model benchmarking, Inference optimization
Preferred skills
LLM serving (vLLM, SGLang, TensorRT-LLM), Inference optimization (quantization, batching, caching), Model fine-tuning, Open-source contributions
Technologies
vLLM, SGLang, TensorRT-LLM, Transformers, Langchain, Langsmith, smolagents, FastAPI, Flask, Kubernetes, Docker, Git, AWS SageMaker, GCP Vertex AI, Azure ML
Responsibilities
Build and test LLM-based solutions and applications; Assist with prompt engineering, model selection, and inference optimization; Run performance and quality experiments for proof-of-concept work; Contribute to internal tooling and automation
Seniority
Early Career (Student/Recent Grad), hands-on IC