Staff Software Engineer, ML Infrastructure
Core
Build and operate large-scale distributed ML infrastructure for real-time computer vision inference and LLM/GenAI serving in a home security platform.
Role type
Staff Software Engineer (ML Infrastructure)
Builds
Kubernetes-based ML platform, real-time CV inference systems, and LLM/GenAI serving infrastructure
Domain
Home Security, Cloud Infrastructure, Distributed Systems, Machine Learning
Deliverable
production ML models | infrastructure
Required skills
distributed systems architecture, Kubernetes, AWS, Kafka, Python, high-throughput low-latency systems, capacity planning, observability, incident response, technical leadership
Preferred skills
Ray, KServe, Triton, vLLM, LLM serving, GPU inference optimization, real-time video pipelines, ML lifecycle management
Technologies
Kubernetes, AWS (EKS, S3, IAM), Kafka, Ray, KServe, Triton, vLLM, Python, Go, C++, Rust
Responsibilities
Drive architecture decisions for the ML platform, build and operate real-time CV inference systems, stand up LLM/GenAI serving infrastructure, mentor engineers, define SLOs and observability standards, lead incident response
Seniority
Staff, hands-on IC with strategic leadership
