Executive Director, ML Engineer MLOps
Core
Build, deploy, and maintain robust distributed training pipelines and high-volume real-time/batch inference systems for large language models and vector databases on GPU-enabled clusters.
Role type
Senior IC MLOps Engineer (LLM inference & distributed training)
Builds
Scalable ML workflows, open-weight LLM serving stacks, and vector database infrastructure for consumer banking personalization.
Domain
Financial services / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Python, AWS, CUDA, quantization (PTQ, AWQ), transformer models, reinforcement learning (GRPO, DAPO), systems engineering (caching, autoscaling, low latency), monitoring & observability, model training & deployment
Preferred skills
recommendation systems, containers (Docker, Kubernetes, ECS), Ray, vLLM, RL libraries (verl, trl), databases
Technologies
AWS, CUDA, Docker, Kubernetes, vLLM, Ray, Python, LLM, MLOps
Responsibilities
Develop and run high-volume real-time and batch inference systems with a focus on performance and reliability; Implement quantization methods and deploy open-weight large language models on modern serving stacks; Oversee the administration and optimization of vector databases; Establish and maintain monitoring and observability pipelines; Collaborate with cross-functional teams to introduce new technologies and improve infrastructure; Partner with product and architecture teams to define scalable technical solutions.