Data Scientist II - Big Data R&D, Identity Graph & Deceased Monitoring
Core
Develop graph-based algorithms and data pipelines on massive PII datasets to power identity verification, entity resolution, and deceased monitoring solutions.
Role type
Mid-level IC data scientist (big data & graph analytics)
Builds
Identity graph, entity-resolution capabilities, and data pipelines for fraud prevention and compliance products
Domain
Identity verification, fraud detection, compliance, big data
Deliverable
production ML models | data pipelines
Required skills
Python, SQL, Spark/PySpark, supervised/unsupervised ML, statistics, data mining, graph techniques
Preferred skills
AWS (EMR, S3), Databricks, Neo4j/AWS Neptune/GraphFrames, Elasticsearch, DynamoDB, Airflow, TensorFlow/PyTorch, XGBoost
Responsibilities
Design and implement ML/data mining/graph algorithms for identity verification; Analyze large datasets to refine entity-resolution and identity-matching algorithms; Build and maintain ETL and feature generation pipelines; Support senior data scientists with feature engineering and error analysis; Evaluate new data sources and design offline experiments; Implement SQL and Python/R code for data extraction and validation; Provide analytical support to compliance teams with ad hoc investigations and dashboards
Seniority
Mid-level, hands-on IC