Senior Site Reliability Engineer for Datacraft team
Core
Senior SRE building the reliability backbone for an AI-first data platform, ensuring deployments, pipelines, and observability for enterprise customers.
Role type
Senior IC Site Reliability Engineer (Data Platform)
Builds
Data-intensive jobs, services, and AI agent infrastructure on GCP and Kubernetes
Domain
Cloud Infrastructure, Data Engineering, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
GCP (BigQuery, DataProc, Cloud Composer, GCS), Kubernetes, Python, Infrastructure as Code (Terraform), CI/CD (GitLab), Observability (OpenTelemetry, Prometheus, Grafana), Incident Management, AI coding agents
Preferred skills
Go, Data pipeline technologies (Kafka, Airflow, Spark, Iceberg), Snowflake, MCP server management
Technologies
GCP, Kubernetes, Terraform, Python, Go, Kafka, Airflow, Cloud Composer, BigQuery, Snowflake, Databricks, Apache Iceberg, GCS, Grafana, Prometheus, Victoria Metrics, PagerDuty, Sentry, OpenTelemetry, GitLab, Cursor, Claude Code
Responsibilities
Build and maintain reliability ecosystem for data services; Ensure end-to-end observability across data pipelines; Automate deployments and operational runbooks; Participate in L3 on-call rotation and incident resolution; Ensure reliability of Loomi Analytics Agent data infrastructure
Seniority
Senior, hands-on IC
