CareerPlanGet AI match score →

Principal Software Engineer, Enterprise Scalability

Boston💼 Full-time💰 $244,000–$244,000🗓 2026-03-31 → 2026-07-31

Core

Principal Software Engineer driving enterprise scalability, performance, and multi-region readiness for Klaviyo's SaaS platform.

Role type

Principal Software Engineer (Enterprise Scalability)

Builds

High-volume, multi-tenant SaaS platform for ecommerce creators

Domain

Ecommerce / SaaS / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Performance engineering, capacity planning, sharding/partitioning, caching/back-pressure, multi-region architecture, high-volume migration, AI-assisted anomaly detection, synthetic load generation, observability

Preferred skills

AI fluency (profiling assistance, workload modeling, runbook copilots), cross-org influence, data-driven bottleneck removal

Technologies

AI tools for scale engineering, synthetic traffic generators, profiling harnesses

Responsibilities

Define enterprise scalability fitness functions and scorecards; design/implement sharding, partitioning, and multi-region strategies; build lightweight enablement (benchmarks, testbeds); lead scalability reviews and readiness gates; integrate AI into scale and resiliency work

Seniority

Principal, hands-on IC with strategic impact

Rewrite
## Responsibilities - Define enterprise scalability fitness functions (latency/throughput/error rates) and a scorecard; align teams to SLOs and budgets. - Design/implement sharding and partitioning strategies, caching/back-pressure, multi-region readiness, and high-volume migration paths. - Build lightweight enablement: benchmarks, profiling harnesses, reproducible testbeds; pair with teams to land fixes. - Lead scalability reviews and readiness gates that accelerate—not block—delivery; drive incident deep dives tied to systemic fixes. - Communicate clearly to execs and engineers, tying technical work to business impact and customer outcomes. - Integrate AI into scale and resiliency work—from proactive anomaly detection to synthetic load and guided runbooks—so performance improvements stick and incidents don’t repeat. ## Requirements - Experience: 12+ years scaling multi-tenant SaaS with a reputation for removing major bottlenecks and proving impact with data. - Technical expertise: Performance engineering, capacity planning, sharding/partitioning, caching/back-pressure, multi-region readiness, and high-volume migrations; you turn hotspots into robust patterns. - AI tools & automation: You apply AI to scale work—profiling assistance, workload modeling, synthetic traffic generation, anomaly detection, and runbook copilots—always with explicit guardrails and observability. - Cross-org influence: You align teams through fitness functions, scorecards, and readiness gates that accelerate—not block—delivery; you communicate tradeoffs crisply to execs and engineers. - AI fluency: Curious, adaptable, and proactive in exploring AI that responsibly improves scale outcomes. ## Nice to Have - Scale scorecard: Company-wide fitness functions (latency/throughput/error rates) are adopted and reviewed regularly. - High-impact wins: 2–3 bottlenecks removed with documented, reproducible testbeds; pXX latencies and error rates improve on top enterprise workloads; repeat P0s trend down. - AI-assisted scale engineering: AI-driven anomaly detection reduces alert noise while improving signal; generative load testing and copilot runbooks are used in release/readiness checks for the top critical services; time-to-isolate regressions drops 20–30%. ## Success in 6–12 Months - Company-wide scale scorecard in place; 2–3 high-impact bottlenecks removed; top enterprise workloads show improved pXX latencies and error rates; fewer repeat P0s.
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗