Inference Infrastructure Architect (Remote)
Core
Design and operate a global, bare-metal inference platform for Telnyx's AI-agent traffic and external customer inference, maximizing throughput per GPU-dollar while meeting latency SLOs.
Role type
Senior/Staff Inference Infrastructure Architect
Builds
Serverless serving pools, fleet layer with disaggregated prefill/decode, dedicated per-tenant pools, and the underlying Kubernetes-on-bare-metal platform.
Domain
Cloud Infrastructure / AI Inference / GPU Computing
Deliverable
production ML models
Required skills
Kubernetes on bare metal, vLLM or SGLang, performance engineering, Python, Go, Linux, networking, storage
Preferred skills
MoE expert parallelism, speculative decoding, KV cache management, GPU scheduling, OpenStack Ironic, Kata Containers
Technologies
vLLM, SGLang, llm-d, NVIDIA Dynamo, Kubernetes, Kueue, Volcano, Mooncake, Dragonfly, Prometheus, OpenTelemetry, DCGM
Responsibilities
Operate and expand the GPU fleet efficiently; implement serverless serving pools and dedicated model deployments; manage weight logistics and elasticity; ensure observability and cost economics; upstream contributions to open-source projects.
Seniority
Senior/Staff, hands-on IC