Software Engineer, Inference - Performance Optimization
Core
Model inference performance across application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference.
Role type
Senior IC systems engineer (inference performance optimization)
Builds
Cost-to-serve estimates, performance models, and tools for latency/capacity/utilization/cost tradeoffs
Domain
AI inference systems, distributed systems, hardware efficiency
Deliverable
production ML models
Required skills
performance profiling, benchmarking, analysis, optimization, distributed systems, model inference, hardware efficiency, systems programming, cross-functional collaboration
Preferred skills
reasoning from first principles, working across abstraction layers (application to kernels/accelerators/networking/fleet scheduling)
Technologies
microbenchmarks, fleet scheduling, accelerators, networking, kernels
Responsibilities
Build and refine performance models translating microbenchmark results into cost-to-serve estimates; Analyze inference workloads end to end across applications, models, and fleet infrastructure; Enhance tooling to identify bottlenecks across layers for latency and throughput; Partner with other teams to turn performance insights into concrete improvements and project future changes
Seniority
Senior, hands-on IC