Sr Engineer, Server Inference
Core
Designing APIs, deploying workloads, and benchmarking end-to-end inference speed for state-of-the-art AI models on Tenstorrent's hardware.
Role type
Senior IC backend engineer (inference server)
Builds
Software layer for AI inferencing on Tenstorrent's cutting-edge hardware
Domain
AI / Machine Learning / Custom Silicon / High-Performance Computing
Deliverable
production ML models
Required skills
API design, system design, performance optimization (batching, caching, model parallelism), Python, Docker, Linux, backend systems architecture
Preferred skills
Clean software architecture, abstraction layers, scaling infrastructure
Technologies
Python, Docker, Linux
Responsibilities
Design modern APIs for ML model deployment, optimize end-to-end ML inference on custom silicon, build scalable and reliable software interfaces for AI workloads, benchmark inference speed
Seniority
Senior, hands-on IC