Software Engineer, Inference AI/ML
Core
Implement well-scoped features and fixes for model-serving services to improve latency, reliability, and cost on a GPU platform.
Role type
IC1 Software Engineer (Inference AI/ML)
Builds
Production features for model-serving services (Triton, vLLM, TensorRT-LLM, Ray Serve)
Domain
Cloud infrastructure for AI inference
Deliverable
production ML models
Required skills
Python, Go, C++, Linux fundamentals, Git/CI, data structures, algorithms, networked services
Preferred skills
PyTorch, TensorFlow, CUDA, Grafana, Prometheus, OpenTelemetry
Technologies
Triton, vLLM, TensorRT-LLM, Ray Serve, Kubernetes
Responsibilities
Implement features and fixes in Python/Go/C++ for model-serving services, Write tests, code comments, and short design docs, Add basic metrics and dashboards, Follow on-call runbooks and learn incident response, Contribute to performance experiments
Seniority
IC1, entry-level with mentorship