Staff Software Engineer, Foundational Model Serving
Core
Design and build high-throughput, low-latency inference systems for hosting and serving frontier AI models (open source and proprietary) via API.
Role type
Staff Software Engineer (Foundational Model Serving)
Builds
Foundation Model Serving API product for hosting and serving frontier AI model inference
Domain
AI Infrastructure / Large Language Model Serving
Deliverable
production ML models
Required skills
large-scale distributed systems, high-scale operationally sensitive backend systems, system design, algorithms, data structures, GPU workload optimization, vLLM, SGLang, token-based rate limiting, architectural trade-offs, code quality, testing, operational readiness, technical mentorship
Preferred skills
product-oriented mindset, cross-functional collaboration, strategic vision
Technologies
vLLM, SGLang, GPU workloads
Responsibilities
Design and implement core systems and APIs ensuring scalability and reliability; Define technical roadmap and long-term architecture; Optimize performance, throughput, and autoscaling for GPU serving workloads; Contribute to key infrastructure components like rate limiters and optimizers; Collaborate cross-functionally to translate customer needs into performant systems; Establish best practices for code quality and mentor engineers; Represent the team in cross-organizational technical discussions
Seniority
Staff, hands-on IC with strategic influence