Senior Software Engineer, Inference
Core
Lead design and optimization of a Kubernetes-native inference platform to improve latency, throughput, and reliability for AI workloads.
Role type
Senior Software Engineer (Inference Systems)
Builds
High-scale AI inference services and orchestration platforms
Domain
Cloud Infrastructure / AI Inference
Deliverable
production ML models
Required skills
Distributed systems, Python, Go, CUDA kernel optimization, Kubernetes, CI/CD, Observability (Prometheus, Grafana, OpenTelemetry), Inference internals (batching, caching, mixed precision, streaming token delivery)
Preferred skills
Inference frameworks (vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe), GPU interconnect topologies (NCCL, SHARP, RDMA, NUMA)
Responsibilities
Lead design reviews and drive architecture; Define and own SLIs/SLOs; Implement advanced optimizations (micro-batch schedulers, speculative decoding, KV-cache reuse); Strengthen incident posture (capacity planning, autoscaling, graceful degradation); Mentor IC1/IC2 engineers; Own an area spanning multiple services and teams
Seniority
Senior, hands-on IC