Senior Software Engineer, Distributed Systems - NIM Factory
Core
Design and build factory infrastructure and automation pipelines to convert AI models into validated, performant NVIDIA Inference Microservices (NIMs) for deployment across heterogeneous cloud, on-prem, and local GPU environments.
Role type
Senior IC distributed systems engineer (factory infrastructure)
Builds
Scalable factory pipelines, containerized microservices, CI/CD automation, and validation infrastructure for NIMs
Domain
AI infrastructure, distributed systems, cloud-native technologies
Deliverable
production ML models
Required skills
Distributed systems architecture, microservices design, containerization (Docker, K8s), cloud technologies, performance engineering, observability, data modeling, schema design, event-driven architecture, CI/CD pipeline development
Preferred skills
Temporal, Kafka, Redis, large-scale full stack development, debugging distributed microservices
Technologies
Docker, Kubernetes, Cloud Endpoints, Helm, Prometheus, Temporal, Kafka, Redis
Responsibilities
Develop factory pipelines to produce deployable services validated across environments; design interfaces and expand observability over compute infrastructure; collaborate with AI model teams to build efficient infrastructure; define metrics and drive improvements based on feedback; mentor team members and grow colleagues
Seniority
Senior, hands-on IC