AI Engineer (LLM Gateway, FM Hosting)
Core
Design, build, and optimize large-scale production AI systems, specifically LLM gateways and foundation model hosting, to improve associate workflows and customer interactions.
Role type
Senior IC AI Engineer (LLM Gateway & Foundation Model Hosting)
Builds
Scalable, responsible AI systems including LLM inference, agentic workflows, similarity search, and multi-model orchestration pipelines.
Domain
Financial Services / Large Language Models / Cloud Infrastructure
Deliverable
production ML models
Required skills
LLM inference optimization, agentic AI system design, multi-model orchestration, cost-performance governance, model right-sizing, vector search, guardrails, cloud platform deployment, hardware utilization optimization, ethical AI standards enforcement
Preferred skills
Foundation model training, mixed AI system architecture, design council leadership, mentorship of senior engineers
Technologies
AWS Ultraclusters, Hugging Face, VectorDBs, PyTorch, CUDA, Golang, Python, Java, C#, Scala
Responsibilities
Design and implement multi-model orchestration pipelines integrating LLMs and domain-specific models; Lead cost-performance governance reviews tracking GPU utilization and inference costs; Mentor Principal and Manager-level AI engineers; Establish standards for ethical AI deployment including explainability and fairness; Optimize training and inference software for hardware utilization and latency; Architect mixed AI systems combining rule-based, retrieval-augmented, and generative components.
Seniority
Senior, hands-on IC with strategic roadmap influence