CareerPlanGet AI match score →

AI Platform Engineer

Bengaluru, India💼 Full-time🗓 2026-05-12 → 2026-07-31

Core

Remove compute-scaling bottlenecks for production LLMs by making frontier-model inference fast, efficient, reliable, and observable.

Role type

Senior IC LLM Inference Engineer (Systems/MLOps)

Builds

Production-grade LLM inference services serving products dependent on frontier models

Domain

Ecommerce, Large Language Models, High-Performance Computing, GPU Systems

Deliverable

production ML models

Required skills

Python, Go or Rust, LLM inference optimization, GPU architecture, CUDA/kernel programming, vLLM, Triton, PyTorch, SGLang, TensorRT, performance tuning, capacity planning, incident response, root-cause analysis

Preferred skills

Quantization, paging, kernel/runtimes improvements, speculative decoding, continuous batching, KV cache management

Technologies

vLLM, Triton, PyTorch, SGLang, TensorRT, CUDA

Responsibilities

Own production inference from handoff to serving including release engineering and incident response; Tune inference performance to reduce latency and increase throughput; Optimize runtimes and servers across heterogeneous GPU fleets; Benchmark and measure latency, throughput, and cost; Improve monitoring, tracing, and alerting for reliability; Apply and ship new optimizations like quantization and kernel improvements; Partner cross-functionally to translate requirements into performance SLOs

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗