Site Reliability Engineer
Core
Improve, manage, and monitor production-critical infrastructure and data pipelines at the intersection of operations and software development.
Role type
Site Reliability Engineer (SRE)
Builds
Proprietary data pipelines and trading systems
Domain
Finance / AI & Machine Learning
Deliverable
infrastructure
Required skills
Python coding and debugging, Linux administration, SQL, incident response, automation, deployment leadership
Preferred skills
gRPC, Postgres, Pandas, Golang, R, Git, Jenkins, Bazel, Prometheus, Grafana, Airflow, Kubernetes, CI/CD best practices
Responsibilities
Improve fault-tolerance and maintainability of code in data pipelines and trading systems; Diagnose and fix bugs in code; Lead complex deployments; Automate manual workflows; Track and prioritize production issues; Share on-call rotation for incident response
Seniority
Mid-level, hands-on IC