Data Engineer, PXT Central Science
Core
Build and maintain scalable data pipelines and production systems to support economics, behavioral science, and machine learning models that improve Amazon's workforce experience.
Role type
Senior Data Engineer (ML Infrastructure)
Builds
End-to-end data engineering solutions, data pipelines, APIs, and data serving layers for production ML models
Domain
Technology / Data Engineering / Machine Learning
Deliverable
production ML models
Required skills
Python, Java, Scala, NodeJS, data modeling, warehousing, ETL pipelines, AWS Redshift, AWS S3, AWS Glue, AWS EMR, AWS Kinesis, AWS FireHose, AWS Lambda, IAM, non-relational databases, Hadoop, Hive, Spark
Preferred skills
Big data technologies (Hadoop, Hive, Spark, EMR)
Technologies
AWS Glue, AWS EMR, AWS Lambda, AWS Redshift, AWS S3, AWS Kinesis, AWS FireHose, Hadoop, Hive, Spark
Responsibilities
Design and maintain scalable data pipelines using native AWS services; develop and maintain APIs and data serving layers for model productionization; build scalable feature extraction and processing frameworks; partner with economics, data science, and software engineering teams to translate analytical requirements into production-ready solutions; maintain layered data systems and build automated reporting solutions
Seniority
Senior, hands-on IC