Junior Data Engineer / 1
Core
Transitioning existing data and ML workflows from batch processing to scalable real-time streaming solutions for ML model inference.
Role type
Junior Data Engineer (Real-time Streaming)
Builds
Scalable, low-latency data pipelines and streaming jobs for ML model inference.
Domain
Data Engineering / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Kafka (producers/consumers, topic design, partitions), Spark Structured Streaming, Kubernetes, Python, ML libraries, batch-to-streaming migration
Preferred skills
Kafka Streams, feature stores, retraining orchestration, ELK/metrics monitoring, performance tuning
Technologies
Kafka, Spark, Kubernetes, Python, MS Teams, Jira, Confluence
Responsibilities
Transform batch inference workflows into streaming pipelines; Define streaming semantics (micro-batching, windowing, state management); Design Kafka topic structures and consumer group patterns; Implement checkpointing and delivery guarantees; Package and version ML model artifacts; Tune performance for throughput and latency; Deploy and operate streaming jobs with monitoring; Integrate streaming outputs into downstream ETL/BI systems; Collaborate on CI/CD for streaming models.
Seniority
Junior, hands-on IC