Sr. Software Engineer, AI / ML Inference Platform
Core
Build and operate the shared AI/ML platform enabling model training, evaluation, and production inference for voice and digital AI agents.
Role type
Senior IC ML Platform Engineer (Inference & Training Systems)
Builds
GPU training infrastructure, model evaluation tooling, and production inference services running on NVIDIA GPUs in GCP.
Domain
AI/ML Infrastructure, Cloud Computing, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Kubernetes, Linux, CI/CD, GPU workload management, distributed systems, model lifecycle management, observability, release automation
Preferred skills
ASR, NLP, LLMs, vLLM, Triton, TGI, model-serving runtimes, performance optimization
Technologies
NVIDIA GPUs, GCP, Kubernetes, Python, Go, vLLM, Triton, TGI
Responsibilities
Design and build shared platform capabilities for model training, evaluation, artifact management, and production inference; optimize GPU training cluster scheduling and resource utilization; develop low-latency, high-throughput inference serving pathways; partner with scientists to translate model capabilities into scalable production designs; build benchmarking and evaluation infrastructure; strengthen telemetry and diagnostic tooling; lead technical projects and mentor engineers.
Seniority
Senior, hands-on IC