AI Engineer 5 (FM Hosting, LLM Inference)
Core
Design, develop, test, deploy, and support large-scale AI systems including foundation model training, LLM inference, multi-agent workflows, and multi-model orchestration pipelines for banking applications.
Role type
Senior IC AI Engineer (LLM Inference & System Optimization)
Builds
Scalable, high-performance AI infrastructure and proprietary solutions empowering teams across Capital One to enhance products with AI.
Domain
Financial Services / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
LLM inference optimization, multi-model orchestration, foundation model training, vector search, system scalability, cost-performance governance, hardware/software optimization, Python, Go, CUDA, Java
Preferred skills
Agentic AI system design, heterogeneous AI system architecture, ethical AI deployment standards, dynamic inference strategies, model compression, right-sizing models and hardware
Technologies
AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, CUDA
Responsibilities
Design and implement multi-model orchestration pipelines integrating LLMs and vector search; Lead cost-performance governance reviews tracking GPU utilization and inference costs; Mentor Principal and Manager-level AI engineers; Contribute to the technical vision and long-term roadmap of foundational AI systems.
Seniority
Senior, hands-on IC with mentorship responsibilities

