Member of Technical Staff
Core
Design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for a generative AI platform.
Role type
Senior IC training infrastructure engineer
Builds
Scalable, resilient backend infrastructure for distributed ML training, inference, and data processing pipelines
Domain
Generative AI infrastructure, cloud-native systems, distributed data systems
Deliverable
infrastructure
Required skills
Large-scale backend infrastructure design, distributed data systems, server-side programming (Python, C++, Go, TypeScript), cloud-native platforms (AWS, GCP, Azure), data processing and API systems, A/B testing and scientific experimentation, cloud-native tools (Docker, Kubernetes)
Preferred skills
Technical design leadership, cross-functional collaboration, operational excellence, data-driven metrics definition
Technologies
Kubernetes, Ray, Kubeflow, MLFlow, PostgreSQL, MySQL, DynamoDB, Apache Spark, Apache Flink, Apache Kafka, gRPC, Thrift, Statsig, Meta Deltoid, Optimizely
Responsibilities
Architect and build scalable, resilient backend infrastructure; Lead technical design discussions and mentor engineers; Design and implement core backend services with focus on efficiency and low latency; Drive infrastructure optimization for compute cost, storage lifecycle, and network performance; Collaborate with ML, DevOps, and product teams; Evaluate and integrate cloud-native and open-source technologies; Own end-to-end systems from design to deployment
Seniority
Senior, hands-on IC with mentorship responsibilities