Senior Site Reliability Engineer
Core
Lead development aspects of the Infrastructure engineering team to design, build, and scale distributed systems for ServiceTitan's cloud platform.
Role type
Senior Site Reliability Engineer (Infrastructure Lead)
Builds
Scalable infrastructure, shared services, and high-availability distributed systems for ServiceTitan's cloud.
Domain
Cloud Infrastructure & Systems Reliability
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Serverless computing, Distributed messaging systems (Kafka, Event Hubs, SQS), Data Lakehouse architectures (Snowflake, Databricks Delta), API gateways, Cloud administration (Azure, AWS), Git, C#, Visual Basic, PowerShell, Java, Infrastructure as Code, CI/CD (Jenkins, Team City), Log/Metric analysis (ELK, DataDog, Grafana)
Preferred skills
None stated
Technologies
Kubernetes, Kafka, Event Hubs, SQS, Snowflake, Databricks Delta, Azure, AWS, Jenkins, Team City, Git, C#, Visual Basic, PowerShell, Java, Elasticsearch, Logstash, Kibana, DataDog, Grafana
Responsibilities
Design, develop, test, troubleshoot, debug, optimize, scale, and maintain software applications; Perform capacity planning and deployment; Collaborate with Product Engineering to plan and deploy releases; Build scalable infrastructure and shared services; Define non-functional requirements for distributed systems; Resolve product/service defects and infrastructure changes.
Seniority
Senior, hands-on IC