Lead Software Engineer - Site Reliability
Core
Design resilient systems, automate recovery, and ensure fast, stable, observable infrastructure at scale for Freshworks' CX and IT solutions.
Role type
Lead Site Reliability Engineer (SRE)
Builds
Automated monitoring, alerting, remediation pipelines, and highly available distributed systems
Domain
SaaS / Customer Experience (CX) & IT Solutions
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Kubernetes orchestration, CI/CD pipeline design, Infrastructure as Code (IaC), distributed system design, observability implementation, disaster recovery strategies, root cause analysis, automation scripting
Preferred skills
Chaos engineering, performance engineering, SRE best practices advocacy
Technologies
Docker, Kubernetes, Linux, IaC tools, monitoring/logging/tracing stacks
Responsibilities
Design tools for availability and scalability; Define SLIs/SLOs and manage error budgets; Build automated monitoring and remediation pipelines; Lead incident response and blameless postmortems; Champion observability across services; Contribute to infrastructure architecture and reliability roadmaps
Seniority
Senior, hands-on IC with leadership responsibilities