Data Platform Engineer (Cloudera)
Core
Design, implement, and maintain distributed computing platforms and Python-based automation/API services for enterprise data reliability and security.
Role type
Senior Data Platform Engineer (Cloudera/Hadoop ecosystem)
Builds
Distributed data processing solutions, secure APIs, and automated services
Domain
Big Data, Cloud Infrastructure, Enterprise Data Platforms
Required skills
Cloudera Hadoop ecosystem (HDFS, YARN, Spark, Impala, Ranger), Python (FastAPI, Gunicorn, spaCy, NumPy, Pandas, Polars), Docker, Kubernetes, Linux/Unix (Bash), Distributed computing principles, Hive, HBase, Scala, PySpark, Cloud services (AWS, Azure, Databricks), Apache Iceberg, Object storage (S3, Ozone)
Preferred skills
Building/managing secure APIs at scale in containerized environments, Agile/Scrum methodologies, Technical documentation
Technologies
Cloudera, Hadoop, Spark, Python, FastAPI, Kubernetes, Docker, AWS, Azure, Databricks, Hive, HBase, Scala, PySpark, Apache Iceberg, S3, Ozone, Bash, Jupyter
Responsibilities
Design and maintain distributed computing platforms; Develop and optimize Python-based automation and API services; Troubleshoot complex issues across data, application, and infrastructure layers; Maintain compliance with enterprise security standards and governance practices; Partner with senior technologists and cross-functional stakeholders to deliver initiatives on schedule
Seniority
Mid-Senior, hands-on IC