Platform Engineer (Machine Learning)
Core
Build and operate the infrastructure and systems powering AI capabilities, from model training and evaluation to deployment, inference, and observability.
Role type
Senior IC ML Platform Engineer
Builds
Production ML systems, model serving infrastructure, and tooling for AI engineers
Domain
AI/ML infrastructure, distributed systems, cloud computing
Deliverable
production ML models | infrastructure
Required skills
Python, PyTorch/JAX, LLM serving (vLLM/SGLang/TensorRT-LLM), distributed systems, GPU infrastructure, vector databases, workflow orchestration
Preferred skills
None stated
Technologies
Python, PyTorch, JAX, vLLM, SGLang, TensorRT-LLM, GPU infrastructure, vector databases
Responsibilities
Design systems for model training, evaluation, deployment, and inference; Build and optimize model serving infrastructure for high-throughput/low-latency; Develop reliable pipelines for data preparation and model release; Build platforms enabling rapid experimentation and model shipping; Develop evaluation and benchmarking infrastructure; Build production observability, monitoring, and alerting; Identify bottlenecks and improve system performance across the ML stack
Seniority
Senior, hands-on IC

