Software Engineer - Model Products
Core
Design, build, and operate Model APIs and inference runtimes to ensure AI models are fast, reliable, and cost-efficient for developers.
Role type
Senior IC backend engineer (LLM inference systems)
Builds
High-performance Model APIs, inference runtimes, and benchmarking frameworks for open-source LLMs
Domain
AI infrastructure / LLM serving / Distributed systems
Deliverable
production ML models
Required skills
distributed systems, large-scale APIs, low-latency backend services, performance profiling, tracing, capacity planning, SLO management, debugging complex systems, API versioning, usage metering, quotas, authentication
Preferred skills
LLM runtimes (vLLM, SGLang, TensorRT-LLM), Kubernetes, service meshes, API gateways, distributed scheduling, open-source APIs
Technologies
TensorRT-LLM, CUDA, JSON mode, grammar-constrained generation, tool/function calling, multi-modal serving, speculative decoding, guided generation, KV-cache reuse
Responsibilities
Design and operate Model APIs with advanced inference capabilities; Profile and optimize TensorRT-LLM kernels and CUDA operators; Productionize performance improvements across runtimes; Build comprehensive benchmarking frameworks; Instrument deep observability and build repeatable benchmarks; Implement platform fundamentals like API versioning and metering; Collaborate with teams to deliver robust model serving experiences
Seniority
Senior, hands-on IC