Senior Site Reliability Engineer for Datacraft team
Core
Senior SRE building and operating the reliability ecosystem for an AI-first data platform, ensuring the stability of data pipelines, orchestration layers, and agentic AI workloads for enterprise customers.
Role type
Senior Site Reliability Engineer (Data Platform)
Builds
Reliable data ingestion pipelines, orchestration services, and agentic AI infrastructure on GCP and Kubernetes.
Domain
Cloud Infrastructure, Data Engineering, AI/ML Operations
Deliverable
production ML models | infrastructure
Required skills
GCP (BigQuery, DataProc, Cloud Composer, GCS), Kubernetes, Python, Infrastructure as Code (Terraform), CI/CD (GitLab), Observability (OpenTelemetry, Prometheus, Grafana), Kafka, Airflow/Cloud Composer, Distributed Tracing, AI Coding Agents
Preferred skills
Go, Apache Spark, Apache Iceberg, MCP server management, LLM API gateway optimization
Responsibilities
Build and maintain reliability ecosystem for data services; Ensure end-to-end observability across data platform; Automate deployments and operational runbooks; Participate in L3 on-call rotation and incident resolution; Ensure reliability of agentic AI platform infrastructure.
Seniority
Senior, independent professional with end-to-end ownership
