CareerPlanGet AI match score →

Software Engineer, Inference - Performance Optimization

San Francisco💼 Full-time🗓 2026-04-25 → 2026-07-31

Core

Model inference performance across application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference.

Role type

Senior IC systems engineer (inference performance optimization)

Builds

Cost-to-serve estimates, performance models, and tools for latency/capacity/utilization/cost tradeoffs

Domain

AI inference systems, distributed systems, hardware efficiency

Deliverable

production ML models

Required skills

performance profiling, benchmarking, analysis, optimization, distributed systems, model inference, hardware efficiency, systems programming, cross-functional collaboration

Preferred skills

reasoning from first principles, working across abstraction layers (application to kernels/accelerators/networking/fleet scheduling)

Technologies

microbenchmarks, fleet scheduling, accelerators, networking, kernels

Responsibilities

Build and refine performance models translating microbenchmark results into cost-to-serve estimates; Analyze inference workloads end to end across applications, models, and fleet infrastructure; Enhance tooling to identify bottlenecks across layers for latency and throughput; Partner with other teams to turn performance insights into concrete improvements and project future changes

Seniority

Senior, hands-on IC

Rewrite
## About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches. ## About the Role In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. ## Responsibilities - Build and refine performance models that translate microbenchmark results into cost-to-serve estimates. - Analyze inference workloads end to end across applications, models, and fleet infrastructure. - Enhance tooling to identify bottlenecks across layers for latency and throughput. - Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference. ## Requirements - Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency. - Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling. - Have deep expertise with performance profiling, benchmarking, analysis, and optimization. - Enjoy collaborating with engineering and research teams to improve real production systems. ## Nice to Have - None specified. ## Benefits - Equal opportunity employer. - Reasonable accommodations available for applicants with disabilities. - Commitment to diversity and inclusion. - Focus on AI safety and human-centric AI development.
Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗