Big Data & Data Infrastructure Engineer
Core
Design and deliver scalable data pipelines and infrastructure for real-time analysis and AI training/inference, partnering with customers and internal teams.
Role type
Senior Solutions Data Engineer (Big Data & Infrastructure)
Builds
Distributed data pipelines, hybrid cloud/on-prem environments, event-driven workflows, and object store-backed data lakes.
Domain
Big Data, Cloud Infrastructure, AI/ML Data Platforms
Deliverable
production ML models
Required skills
Python, Apache Spark (batch & streaming), Apache Kafka, Trino, Terraform, Kubernetes, Docker, SQL, NoSQL, HDFS, distributed systems, stream processing, performance benchmarking (via careerplan.io/jobs/43-001-BF-546-big-data-data-infrastructure-engineer-at-vast-data)
Preferred skills
Bash scripting, low-level debugging, high-level architecture design, technical documentation, cross-functional collaboration
Technologies
Kafka, Spark, Python, Trino, Airflow, S3, Terraform, Docker, Kubernetes, CI/CD, Parquet, ORC
Responsibilities
Build distributed data pipelines using Kafka, Spark, and S3-compatible data lakes; Design and troubleshoot hybrid cloud/on-prem environments; Implement event-driven and serverless workflows; Create technical guides and architecture docs; Integrate data validation and observability tools; Own end-to-end platform lifecycle from ingestion to compute; Benchmark and tune storage and compute layers; Work cross-functionally with R&D on performance limits.
Seniority
Mid-Senior, hands-on IC with customer-facing responsibilities
