Site Reliability Engineer
Core
Ensuring the reliability, availability, and performance of a secondary market trading platform for venture-backed companies, including supporting AI/ML workloads.
Role type
Site Reliability Engineer (Infrastructure & AI Systems)
Builds
Scalable, resilient infrastructure for trading platform and AI/ML model-serving systems
Domain
Fintech / Secondary Markets / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Elixir, Kubernetes, Terraform, AWS (EKS, RDS, VPC), PostgreSQL, Datadog
Preferred skills
Regulated environment experience, CI/CD (GitHub Actions), SOC 2 compliance, Cloudflare, AI/ML production support
Technologies
Elixir, Kubernetes, Terraform, AWS, Vercel, PostgreSQL, Datadog, Cloudflare, GitHub Actions
Responsibilities
Maintain platform uptime and availability; optimize infrastructure for reliability and security; proactively resolve scaling issues; partner with product engineers on performance troubleshooting; configure monitoring and observability systems; lead incident response and postmortems; support and scale infrastructure for AI/ML systems; improve observability for AI workloads
Seniority
Mid-level to Senior, hands-on IC
