Senior Lead AI Engineer (FM Hosting, LLM Inference)
Core
Design, develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, and observability to build scalable, high-performance AI infrastructure.
Role type
Senior Lead AI Engineer (LLM Inference & Hosting)
Builds
Proprietary AI systems and platforms that empower teams across the bank to enhance products with transformative AI capabilities.
Domain
Banking / Applied AI / Large Language Models
Deliverable
production ML models
Required skills
LLM inference optimization, foundation model training, similarity search, vector databases, model evaluation, system observability, hardware/software optimization, Python, Go, Scala, Java
Preferred skills
Cloud platform deployment (AWS/GCP/Azure), leading and mentoring engineering teams, cross-functional stakeholder influence, C++/C#/Golang development, state-of-the-art training/inference optimization techniques
Technologies
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch
Responsibilities
Partner with cross-functional teams to deliver AI-powered products; Invent and introduce state-of-the-art LLM optimization techniques to improve scalability, cost, latency, and throughput; Contribute to the technical vision and long-term roadmap of foundational AI systems.
Seniority
Senior Lead, hands-on IC with mentorship responsibilities

