Site Reliability Engineer
Core
Ensuring the reliability, availability, and performance of a secondary market trading platform for venture-backed companies, including supporting AI/ML workloads.
Role type
Site Reliability Engineer
Builds
Scalable and resilient infrastructure for trading platform and AI/ML systems
Domain
Fintech / Secondary Market Trading / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Elixir, Kubernetes, Terraform, AWS (EKS, RDS, VPC), PostgreSQL, Datadog
Preferred skills
CI/CD, SOC 2 compliance, Cloudflare, AI/ML model serving, vector databases
Technologies
Elixir, Kubernetes, Terraform, AWS, Vercel, PostgreSQL, Datadog, GitHub Actions, Cloudflare
Responsibilities
Maintain platform uptime and availability; optimize infrastructure for reliability and security; proactively resolve scaling issues; partner with engineers on performance troubleshooting; configure monitoring and observability; lead incident response and postmortems; support and scale AI/ML infrastructure; improve observability for AI systems
Seniority
Mid-level, hands-on IC
