CareerPlanSign in

Software Engineer, Inference - Performance Optimization

San Francisco💼 Full-time🗓 2026-04-25 → 2026-09-26

Core

Model inference performance across application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference.

Role type

Senior IC systems engineer (inference performance optimization)

Builds

Cost-to-serve estimates, performance models, and tools for latency/capacity/utilization/cost tradeoffs

Domain

AI inference systems, distributed systems, hardware efficiency

Deliverable

production ML models

Required skills

performance profiling, benchmarking, analysis, optimization, distributed systems, model inference, hardware efficiency, systems programming, cross-functional collaboration

Preferred skills

reasoning from first principles, working across abstraction layers (application to kernels/accelerators/networking/fleet scheduling)

Technologies

microbenchmarks, fleet scheduling, accelerators, networking, kernels

Responsibilities

Build and refine performance models translating microbenchmark results into cost-to-serve estimates; Analyze inference workloads end to end across applications, models, and fleet infrastructure; Enhance tooling to identify bottlenecks across layers for latency and throughput; Partner with other teams to turn performance insights into concrete improvements and project future changes

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.