CareerPlanGet AI match score →

Manager, Software Engineering - Observability

Canada🌐 Remote💼 Full-time💰 $258,000–$258,000🗓 2026-04-15 → 2026-07-31

Core

Lead a team of engineers building and operating Figma's observability stack to ensure platform health, performance, and cost efficiency.

Role type

Senior Engineering Manager (Platform/Observability)

Builds

Core observability stack, instrumentation libraries, agents, operators, and cost attribution frameworks for Figma's platform.

Domain

SaaS / Distributed Systems / Observability

Deliverable

production ML models | infrastructure

Required skills

Team leadership, distributed systems architecture, observability platform strategy, cost optimization, AI-driven anomaly detection, cross-functional alignment, technical mentorship

Preferred skills

Company-wide observability standards design, vendor negotiation, machine learning for operations, multi-language instrumentation, scaling engineering teams

Technologies

Datadog, OpenTelemetry, metrics, logs, distributed tracing, cloud environments

Responsibilities

Lead and grow a team of engineers; define technical strategy for instrumentation and cost transparency; implement AI-driven anomaly detection; establish cost attribution and budgeting frameworks; partner with infrastructure, product, finance, and security teams; optimize observability footprint and spend; coach and mentor engineers

Seniority

Senior, hands-on IC leader

Rewrite
## Responsibilities - Lead and grow a team of engineers responsible for the reliability, scalability, and evolution of Figma’s observability and cost engineering platforms - Own and operate Figma’s core observability stack, including vendor platforms such as Datadog, ensuring high availability, strong data quality, and effective signal-to-noise across metrics, logs, and traces - Define and drive the technical strategy for instrumentation standards, observability libraries, agents, and operators used to monitor internal and external facing services - Explore and implement innovative, AI-driven approaches to anomaly detection, root cause analysis, signal correlation, and operational automation - Establish clear frameworks for cost attribution, budgeting, forecasting, and alerting across infrastructure and observability spend, enabling teams to make informed tradeoffs - Partner with infrastructure, product engineering, finance, and security teams to improve visibility into system health and cost efficiency at scale - Lead initiatives to optimize observability footprint and spend, balancing depth of insight with performance and cost considerations - Coach and mentor engineers through career development, performance feedback, and technical leadership, fostering a culture of ownership, collaboration, and high quality execution ## Requirements - 4+ years of experience leading infrastructure, observability, or platform engineering teams, with a track record of delivering highly reliable production systems - Deep hands-on experience with modern observability platforms (e.g., Datadog, OpenTelemetry) across metrics, logs, and distributed tracing - Strong understanding of distributed systems, instrumentation best practices, SLO design, and incident response workflows - Experience driving cost transparency and accountability initiatives, including cost attribution, budgeting, forecasting, and alerting in cloud environments - Demonstrated ability to set technical direction, drive cross-functional alignment (Engineering, Finance, Security), and make sound architectural decisions in complex environments ## Nice to Have - Experience designing or evolving company-wide observability standards, shared libraries, and agent/operator-based integrations - Background in cost optimization for infrastructure or observability tooling, including vendor negotiations and usage modeling - Experience applying AI or machine learning techniques to anomaly detection, root cause analysis, or operational automation - Familiarity with OpenTelemetry and modern instrumentation frameworks across multiple programming languages - Experience scaling and mentoring high-performing engineering teams through platform expansion or significant architectural change ## Benefits At Figma, one of our values is Grow as you go. We believe in hiring smart, curious people who are excited to learn and develop their skills. If you’re excited about this role but your past experience doesn’t align perfectly with the points outlined in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗