Staff Software Engineer, Foundational Model Serving
Core
Design and build high-throughput, low-latency inference systems for hosting and serving frontier AI models (open source and proprietary) via API.
Role type
Staff Software Engineer (Foundational Model Serving)
Builds
Foundation Model Serving API product for LLM inference
Domain
AI Infrastructure / Large Language Model Serving
Deliverable
production ML models
Required skills
Large-scale distributed systems, high-scale operational backend systems, system design, algorithms, data structures, GPU workload optimization, architectural decision-making, code quality, testing, operational readiness, technical mentorship
Preferred skills
Experience with vLLM or SGLang, token-based rate limiting, cross-functional collaboration, product-oriented mindset
Technologies
vLLM, SGLang, GPU workloads
Responsibilities
Design and implement core systems and APIs ensuring scalability and reliability; Define technical roadmap and long-term architecture; Optimize performance, throughput, and autoscaling for GPU serving; Contribute to key infrastructure components; Collaborate cross-functionally to translate customer needs into systems; Establish best practices and mentor engineers; Represent the team in cross-organizational technical discussions
Seniority
Staff, hands-on IC with strategic influence