Director, AI Platform Reliability
Core
Lead the architecture, development, and operation of highly scalable, distributed software platforms processing hundreds of millions of transactions and terabytes of data.
Role type
Director, AI Platform Reliability (Strategic leadership with hands-on technical depth)
Builds
High-volume, business-critical distributed systems, data platforms, and cloud-native microservices
Domain
IT Infrastructure Management, Cloud-Native, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Java, JVM performance, Apache Kafka, distributed systems, microservices, Kubernetes, cloud platforms (AWS/GCP/Azure), observability, capacity planning, engineering leadership
Preferred skills
Data lakehouse architecture, CI/CD automation, technical debt management, disaster recovery design
Technologies
Java, Kafka, Kubernetes, AWS, Google Cloud, Microsoft Azure
Responsibilities
Define technical strategy and architecture for high-volume data platforms; Guide development of Java-based microservices and streaming pipelines; Establish reliable data ingestion, transformation, and governance practices; Build low-latency, fault-tolerant systems; Define and own operational SLAs/SLOs; Drive capacity planning and latency optimization; Partner with cross-functional teams to deliver platform initiatives; Recruit and mentor engineering leaders.
Seniority
Director, strategic leadership with hands-on IC contribution