Senior Software Engineer - Reliability, Infrastructure, and Tooling
Core
Building reliable, scalable infrastructure for demanding real-time and AI workloads, enabling engineering teams to self-serve capabilities without operational bottlenecks.
Role type
Senior IC infrastructure and reliability engineer
Builds
Production infrastructure, developer tooling, and observability systems for global-scale real-time and AI applications
Domain
Cloud infrastructure, distributed systems, real-time media, and secure execution environments
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Linux internals, networking, distributed systems, observability, incident management, automation, system design
Preferred skills
Kafka, ClickHouse, global Layer 3 networking, real-time media operations, Google SRE experience, PCI compliance
Technologies
Kubernetes, Kafka, ClickHouse, Linux
Responsibilities
Design and ship reliability-focused engineering work within production codebases; build and evolve internal infrastructure and developer tooling; develop observability capabilities; participate in on-call rotation and incident response; investigate complex system-level problems; improve configuration management and reduce technical debt; contribute to architectural discussions; automate repetitive operational processes.
Seniority
Senior, hands-on IC
