Customer Reliability Engineer (CRE)
Core
Stabilize and automate mission-critical customer-facing network infrastructure operations, transitioning from manual incident response to programmatic solutions.
Role type
Mid-level Customer Reliability Engineer (SRE)
Builds
Automated tooling, declarative infrastructure, and observability pipelines for network detection and response platforms.
Domain
Data center networking, cloud infrastructure, and security operations.
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux systems administration, TCP/IP and DNS networking fundamentals, AWS (VPC, EC2, IAM, S3), Terraform, CI/CD pipelines, Python or Go, Bash scripting, incident response, root-cause analysis, SLO definition, on-call management, airgapped environment deployment.
Preferred skills
Experience with Scala, C, C++, Haskell, or Rust, complex log filtering, system instrumentation.
Technologies
Python, Go, Bash, Terraform, AWS, Linux, CloudVision, EOS.
Responsibilities
Perform manual deployments and reactive troubleshooting for critical customer systems, lead incident response and stabilization, write code to automate operational tasks and eliminate manual toil, build and manage observability stacks for metrics, logs, and traces, collaborate with product engineering to design durable system-level fixes, conduct on-site deployments for airgapped customers.
Seniority
Mid-level, hands-on IC