Sr. Software Engineer - AI Inference & Serving
Core
Design and implement core inference infrastructure for serving frontier AI models (LLMs, GenAI) in production, focusing on performance, efficiency, and scaling.
Role type
Senior IC software engineer (AI inference & serving)
Builds
High-availability inference platform for state-of-the-art models (OpenAI, Anthropic, xAI) including GPT5, Realtime audio, and Sora.
Domain
Generative AI, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
C, C++, C#, Java, distributed computing, architecture, high-scale online systems, low latency, high throughput, L7 network proxies, HTTP, TCP, Docker, Kubernetes, Golang
Technologies
C, C++, C#, Java, Docker, Kubernetes, Golang
Responsibilities
Design efficient load scheduling and balancing strategies; Scale the platform to support growing inferencing demand; Collaborate with internal and external partners.
Seniority
Senior, hands-on IC