Senior Site Reliability Engineer
Core
Build and scale internal platform offerings (compute, storage, networking) to ensure reliability and performance of applications.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Internal compute, storage, and networking services for a global private market infrastructure platform
Domain
Cloud Infrastructure / Private Markets
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Cloud platforms (AWS/GCP/Azure), Kubernetes, Infrastructure as Code (Terraform/Ansible), Networking (CNI/Service Mesh), Monitoring (Prometheus/Grafana/ELK/Datadog), Python, API design (REST/GraphQL), CI/CD
Preferred skills
Service mesh implementation, AI tool usage for automation
Technologies
Python, Java, Terraform, gRPC, Docker, Kubernetes, Postgres, AWS, Prometheus, Grafana, ELK Stack, Datadog
Responsibilities
Design and implement monitoring, alerting, and incident response systems; Collaborate with application engineers to guide scalable design; Operate CI/CD pipelines; Push boundaries to incrementally improve systems
Seniority
Senior, hands-on IC
