CareerPlanGet AI match score →

Distributed AI Support Engineer

Athens, Attica, Greece💼 Full-time🗓 2026-01-13 → 2026-07-31

Core

Provide first-line support, debugging, and operations for AI/LLM workloads on high-performance computing (HPC) infrastructure, enabling researchers and industry teams to train and run large-scale models.

Role type

Distributed AI Support Engineer (IC)

Builds

AI/LLM workflows, containerized environments, and support documentation for HPC and cloud stacks.

Domain

High-Performance Computing (HPC) and Artificial Intelligence

Deliverable

production ML models | infrastructure

Required skills

Python programming, PyTorch, distributed training frameworks (DDP, FSDP, DeepSpeed), GPU debugging, container management (Apptainer/Singularity), HPC job scheduling (Slurm), LLM inference frameworks (vLLM, Ray)

Preferred skills

LLM fine-tuning and quantization (QLoRA, PEFT), profiling tools (NVIDIA Nsight, TensorBoard), data I/O optimization, user training and documentation

Technologies

PyTorch, TensorFlow, Hugging Face Transformers, DeepSpeed, vLLM, Ray, Slurm, Apptainer, NVIDIA CUDA, NCCL, PyTorch Profiler, MLflow, Weights & Biases

Responsibilities

Triage and diagnose AI/HPC job failures, support users in writing and debugging multi-GPU job scripts, maintain and test shared AI/LLM software stacks, profile and tune distributed training workloads, advise on data storage and I/O bottlenecks, monitor usage metrics and prepare technical reports, develop documentation and training materials

Seniority

Mid-level, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workable ↗