Site Reliability Engineer
Core
Improve, manage, and monitor production-critical infrastructure and data pipelines for AI/ML investment management systems.
Role type
Site Reliability Engineer (SRE)
Builds
Proprietary data pipelines and trading systems
Domain
Finance / AI & Machine Learning
Deliverable
infrastructure
Required skills
Python coding and debugging, Linux administration, SQL, incident response, automation, deployment leadership
Preferred skills
gRPC microservices, Postgres, Pandas, Golang, R, Git, Jenkins, Bazel, Prometheus, Grafana, Airflow, Kubernetes, CI/CD best practices
Responsibilities
Improve fault-tolerance and maintainability of code in proprietary data pipelines and trading systems; Diagnose and fix bugs in code; Lead complex deployments; Automate manual workflows; Track and prioritize outstanding production-related issues; Share an on-call rotation responding to incidents to ensure the continuous operation of production-critical systems
Seniority
Mid-level, hands-on IC