ML/AI Engineer
Core
Build, automate, and maintain scalable systems supporting the full machine learning lifecycle, including Kubernetes orchestration, CI/CD automation, GPU optimization, and large-scale model deployment.
Role type
Senior IC ML/AI Engineer (MLOps & Infrastructure)
Builds
Production-grade ML inference and training services, CI/CD pipelines, and observability systems for financial services.
Domain
Financial Services / MLOps / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Kubernetes, Docker, Helm, CI/CD (Harness), GitOps, CUDA, TensorRT, NVIDIA Triton, TorchServe, Prometheus, Grafana, Dynatrace, MLflow
Preferred skills
GCP (GKE, Vertex AI), LangChain/LangGraph, Ray, Kubeflow, RLHF workflows, Model Context Protocol (MCP)
Technologies
Kubernetes, Harness, Git, Prometheus, Grafana, Dynatrace, MLflow, NVIDIA Triton, TorchServe, CUDA, TensorRT, GCP, Vertex AI, LangChain, Ray, Kubeflow
Responsibilities
Compose and operate production-grade Kubernetes clusters for high-volume model inference and scheduled training jobs; Configure autoscaling, resource quotas, GPU/CPU node pools, service mesh, and Helm charts; Implement GitOps workflows for environment configuration and application releases; Build CI/CD pipelines to automate build, test, model packaging, and deployment; Enable progressive delivery strategies and integrate quality gates; Standardize pipelines for continuous training and monitoring; Deploy and tune GPU-backed inference services; Implement end-to-end observability for models and pipelines; Establish actionable alerting and runbooks for on-call operations; Operate a model registry with experiment tracking and versioning; Enforce audit readiness with model cards and reproducible builds.
Seniority
Senior, hands-on IC