Generative AI Systems Engineer – Vision-Language Models
Core
Design, evaluate, and optimize Vision-Language Model (VLM) systems for real-world applications, focusing on tradeoff analysis across accuracy, latency, and cost.
Role type
Generative AI Systems Engineer (Vision-Language Models)
Builds
Scalable inference pipelines, serving layers (API/service), and data processing pipelines for multimodal models.
Domain
Generative AI, Multimodal Machine Learning, Computer Vision, NLP
Deliverable
production ML models
Required skills
Model evaluation and benchmarking, parameter-efficient fine-tuning (LoRA, QLoRA), ML system design (inference, scaling, optimization), data pipeline construction, Python, PyTorch
Preferred skills
Experience with VLM architectures (LLaVA, BLIP, Flamingo), model optimization techniques (quantization, batching, caching), Docker/containerized deployments, large-scale dataset handling
Technologies
PyTorch, LoRA, QLoRA, Docker
Responsibilities
Evaluate pretrained VLMs on domain-specific datasets and define evaluation metrics; Implement parameter-efficient fine-tuning techniques; Design controlled experiments to compare baseline vs improved models; Architect scalable inference pipelines for multimodal models; Build pipelines to process and align images, textual queries, and structured metadata.
Seniority
Mid-Senior (5–7 years experience)