Software Engineer 4/5 – Model Serving Systems, AI Platform
Core
Build and scale compute infrastructure for real-time AI/ML model inference and serving, including support for LLMs and large foundation models.
Role type
Senior IC software engineer (model serving systems)
Builds
Scalable, high-availability model serving platforms and infrastructure for Netflix's consumer and studio-facing AI/ML applications.
Domain
Internet streaming, AI/ML infrastructure, large language models
Deliverable
production ML models
Required skills
distributed systems engineering, high-traffic service architecture, object-oriented programming (Java), performance tuning, deployment management, capacity planning, observability, logging, cloud platform expertise (AWS/Azure/GCP), ML model deployment tools (Triton, TensorRT, Docker)
Preferred skills
experience with generative models and LLMs, reducing inference latency and costs, streamlining research-to-production workflows
Technologies
Java, Triton Inference Server, TensorRT, Docker, AWS, Azure, GCP
Responsibilities
Develop and expand compute infrastructure to support growing AI needs, optimize model serving for high availability and performance, partner with ML engineers and data scientists to enable new business areas, promote best practices in observability and logging
Seniority
Senior, hands-on IC