Senior Technology Site Reliability Engineer
Core
Ensuring the reliability, scalability, and performance of critical infrastructure and applications by blending software and systems engineering to build automated, resilient, and observable systems.
Role type
Senior IC Site Reliability Engineer
Builds
Automated, resilient, and observable systems supporting high availability and operational excellence
Domain
Legal technology infrastructure, cloud platforms, and distributed systems
Deliverable
production ML models | infrastructure
Required skills
Terraform, Python/Go/Java, AWS, container orchestration, distributed systems, performance tuning, automation, configuration management (Puppet/Chef/Salt)
Preferred skills
ETL data workflows (AWS EMR, Azure Synapse, Airflow), Kubernetes (AKS/EKS/GKE), Data Lake environments (DataBricks, Snowflake)
Technologies
Terraform, Prometheus, Grafana, DataDog, AWS, Kubernetes, Puppet, Chef, Salt, Python, Go, Java, Apache Hive, Spark, Airflow, Azure Synapse, Azure Data Factory, DataBricks, Snowflake
Responsibilities
Monitor and maintain production systems for high availability; Implement and manage SLIs, SLOs, SLAs, and error budgets; Participate in on-call rotations and incident response with root cause analysis; Develop and maintain infrastructure as code (IaC); Automate deployment, scaling, and recovery processes; Partner with DevOps to build and maintain CI/CD pipelines; Implement observability solutions using metrics, logs, and traces; Proactively identify and resolve system bottlenecks and reliability risks; Document operational procedures and share knowledge across teams
Seniority
Senior, hands-on IC