Inference Engineer
Core
Build end-to-end inference capabilities on a unified control plane (Forge) to deploy and serve AI models across distributed, heterogeneous hardware clusters.
Role type
Senior Inference Engineer (Infrastructure)
Builds
Production-ready model serving infrastructure, monitoring, gateways, and endpoints for a global GPU marketplace.
Domain
AI Infrastructure / GPU Cloud / Distributed Systems
Deliverable
production ML models
Required skills
Kubernetes cluster operations, inference frameworks evaluation, NVIDIA Dynamo architecture, KV-cache orchestration, autoscaling, production monitoring setup, end-to-end product building
Preferred skills
Model optimization (quantization, batching, kernel tuning), RDMA networking, heterogeneous accelerator deployment, customer debugging support
Technologies
Kubernetes, NVIDIA Dynamo, Forge, NeoCloud
Responsibilities
Deploy and serve models on heterogeneous hardware clusters, evaluate and select inference frameworks, build monitoring and gateways for production readiness, optimize inference performance and manage autoscaling, orchestrate KV-cache, debug customer inference issues
Seniority
Senior, hands-on IC