Machine Learning Infrastructure Engineer, Model Inference
Core
Design, deploy, and maintain scalable Kubernetes clusters and model serving infrastructure for real-time AI model inference and training in healthcare.
Role type
Senior IC ML Infrastructure Engineer (Model Inference)
Builds
Scalable Kubernetes clusters, high-performance model serving infrastructure, and robust model API orchestration systems.
Domain
Healthcare + Distributed Systems / Cloud Infrastructure
Deliverable
production ML models
Required skills
Kubernetes administration, distributed systems architecture, API development, GPU cluster management, compute-heavy workflow optimization
Preferred skills
NVIDIA Triton Server, VLLM, TRT-LLM, PyTorch, Tensorflow, CUDA optimization, Infrastructure as Code (Terraform, Ansible), GitOps, container registry management
Technologies
Kubernetes, NVIDIA Triton Server, VLLM, TRT-LLM, PyTorch, Tensorflow, Terraform, Ansible, CUDA
Responsibilities
Design and maintain scalable Kubernetes clusters for AI model inference and training; Develop and optimize ML model serving infrastructure for high-performance and low-latency; Collaborate with teams to scale backend infrastructure for model deployment and throughput optimization; Optimize compute-heavy workflows and enhance GPU utilization; Build a robust model API orchestration system; Define and implement strategies for scaling infrastructure as the company grows.
Seniority
Senior, hands-on IC