CareerPlanGet AI match score →

Sr Engineer, Server Inference

Belgrade💼 Full-time🗓 2026-07-11 → 2026-07-31

Core

Designing APIs, deploying workloads, and benchmarking end-to-end inference speed for state-of-the-art AI models on Tenstorrent's hardware.

Role type

Senior IC backend engineer (inference server)

Builds

Software layer for AI inferencing on Tenstorrent's cutting-edge hardware

Domain

AI / Machine Learning / Custom Silicon / High-Performance Computing

Deliverable

production ML models

Required skills

API design, system design, performance optimization (batching, caching, model parallelism), Python, Docker, Linux, backend systems architecture

Preferred skills

Clean software architecture, abstraction layers, scaling infrastructure

Technologies

Python, Docker, Linux

Responsibilities

Design modern APIs for ML model deployment, optimize end-to-end ML inference on custom silicon, build scalable and reliable software interfaces for AI workloads, benchmark inference speed

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗