Head of Data Platform
Core
Designing the data substrate for a Dealer Intelligence Platform that turns a transactional automotive SaaS platform into a decision engine by building the lake, feature store, model registry, and action ledger.
Role type
Head of Data Platform
Builds
A closed-loop Dealer Intelligence Platform serving thousands of dealer groups with optimization functions for pricing, lead routing, inventory, and scheduling.
Domain
Automotive SaaS / Data Engineering / Machine Learning Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Multi-tenant data architecture, Feature Store design, Data Lake governance, Real-time data streaming, Data migration strategy, Data quality monitoring, Model serving infrastructure, Cloud data platforms (AWS), Optimization loop design, Schema evolution
Preferred skills
Experience with high-throughput transactional platforms, Knowledge of DMS/CRM integrations, Expertise in action telemetry and feedback loops
Technologies
MySQL, Aurora, DynamoDB, S3, Glue Data Catalog, SageMaker, Bedrock, Lake Formation, SOAP/XML, REST, JSON
Responsibilities
Own the end-to-end data platform including Bronze/Silver/Gold layers, Feature Store architecture, Action Ledger implementation, Model data pipeline, Data migration strategy, and Data quality/anomaly detection.
Seniority
Senior, hands-on IC with strategic ownership
Full job description
We operate a multi-tenant automotive SaaS platform serving thousands of dealer groups across the United States. Our data layer — MySQL, Aurora, DynamoDB with DynamoDB Streams, S3, Glue Data Catalog — has grown to support a complex, high-throughput transactional platform. That layer works. Now we need to make it intelligent. We are building a Dealer Intelligence Platform: a closed-loop system that observes raw signals from dealer operations, predicts outcomes, optimizes decisions under constraints, acts through approved channels, and learns from what happened. Pricing optimization, lead routing, inventory mix planning, service bay scheduling — each is a self-contained optimization function that consumes features, scores predictions, and closes the loop with action telemetry. This role owns the entire data substrate that makes that loop possible — the lake, the feature store, the model registry, the action ledger, and the governance framework that keeps it all tenant-isolated and audit-grade. You are not inheriting a finished architecture; you are designing the one that turns a transactional platform into a decision engine.
Scope & Scale
Targeting 5,000+ dealer tenants, each with isolated databases and per-tenant configuration. Billions in annual Gross Merchandise Value (GMV) flowing through platform transactions. 30+ third-party integrations across DMS, CRM, lending, F&I, and marketplace providers — each pushing data in different formats (SOAP/XML, REST, JSON, email). Data pipelines spanning 6 integration domains with multi-protocol vendor connectivity. Data Streams processing real-time change events across onboarding, inventory, and transaction tables. A Product roadmap with 6+ optimization functions — each requiring its own entity model, feature set, constraint definition, and feedback path.
What You Will Own
The data platform end to end: Bronze (raw + telemetry), Silver (canonical entities via Standardization Agent), Gold (KPIs + derived features) — plus the Feature Store and Action Ledger that make the optimization loop possible.
Feature Store architecture — online (sub-50ms reads for real-time scoring) and offline (point-in-time joins for training). Feature contract: owner, freshness SLO, PII tag, training/inference parity. Governed by Lake Formation.
Action Ledger — every recommendation, approval, override, and outcome logged as a first-class object. This is the substrate that closes the loop: without it, models cannot retrain, we cannot attribute lift, and we cannot prove value to dealers.
Model data pipeline — the feature materialization, training data assembly, and serving infrastructure that feeds SageMaker models and Bedrock agents across all optimization functions.
Data migration strategy for the legacy platform — defining which tables move to DynamoDB, which consolidate into Aurora Serverless, and how dual-write validation works at every stage.
Data quality and anomaly detection — automated monitoring for schema d