CareerPlanSign in

Data Engineer II - Getting Customers Ready for AI

United States, Washington, Redmond💼 Full-time🗓 2026-06-22 → 2026-09-26

Core

Design and build scalable batch and streaming data pipelines to process security and operational telemetry, transforming raw data into structured datasets for AI/ML workloads.

Role type

Data Engineer II

Builds

Scalable data pipelines, ETL/ELT workflows, data ingestion frameworks, and structured datasets for AI/ML teams.

Domain

Cybersecurity / Data Engineering / AI Infrastructure

Deliverable

production ML models

Required skills

Python, Scala, SQL, Spark, Flink, Kafka, Event Hub, Azure ADLS, Azure Synapse, Azure Databricks, data modeling, ETL processes, distributed data systems, data governance, data validation, observability, system design.

Preferred skills

C, C++, C#, Java, JavaScript, feature stores, training data preparation, RAG pipelines, telemetry metrics systems.

Responsibilities

Design and build scalable data pipelines (batch and streaming) to process large volumes of security and operational data; Develop and optimize ETL/ELT workflows that transform raw telemetry into structured, consumable datasets; Implement data ingestion frameworks to integrate multi-source data from services, APIs, and event streams; Improve pipeline performance, reliability, and efficiency through partitioning, indexing, and optimization techniques; Design and evolve data models, schemas, and storage strategies for analytics and AI use cases; Ensure data is structured for downstream ML pipelines, feature engineering, and analytics workloads; Implement data validation, quality checks, and observability frameworks to ensure data accuracy and reliability; Apply best practices for data governance, lineage, and auditing across pipelines and datasets; Ensure compliance with security, privacy, and regulatory requirements when handling sensitive data; Contribute to standardization of data contracts and schemas across services; Enable high-quality datasets for AI/ML teams, including support for feature pipelines and training data preparation; Collaborate with AI engineers to ensure data is optimized for RAG pipelines, model training, and evaluation workflows; Build and maintain telemetry pipelines and metrics systems that provide insights into AI system performance and usage; Support development of data-driven signals and insights that improve customer security posture; Contribute to system design discussions, architecture reviews, and cross-team integration efforts; Work across teams to ensure consistent data definitions and interoperability across platforms; Build monitoring and alerting mechanisms for pipeline health, latency, and data quality issues.

Seniority

Mid-level, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.