Research Engineer, Privacy and Anonymization
Core
Build systems to detect and remove sensitive information (PII, credentials) from raw data to protect privacy while preserving utility for AI training and synthetic data workflows.
Role type
Research Engineer (Privacy and Anonymization)
Builds
Production pipelines for data anonymization, detection systems, and evaluation frameworks for privacy risk and data utility.
Domain
AI data infrastructure, privacy engineering, synthetic data generation
Deliverable
production ML models | infrastructure
Required skills
Python, information extraction, named-entity recognition, classification, LLM-based methods, data processing pipelines, redaction, masking, pseudonymization, anonymization, synthetic data generation, schema drift handling, low-latency ML inference
Preferred skills
differential privacy, k-anonymity, secure aggregation, format-preserving encryption, high-throughput data processing, healthcare/finance/security domain experience
Technologies
Python, LLMs
Responsibilities
Build systems to detect PII and sensitive information; develop and benchmark detection approaches combining rules, statistical models, and LLMs; create evaluation frameworks for privacy risk and data utility; design robust systems for schema drift and edge cases; collaborate to translate privacy requirements into technical policies.
Seniority
Mid-level, hands-on IC
