Member of Technical Staff
Core
Design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for a generative AI platform.
Role type
Senior IC training infrastructure engineer
Builds
Scalable, resilient backend infrastructure for distributed ML training, inference, and data processing
Domain
Cloud-native infrastructure for generative AI and distributed data systems
Deliverable
infrastructure
Required skills
large-scale backend infrastructure design, distributed data systems, cloud-native platforms, server-side programming, technical design documentation, cross-functional project leadership, data processing and API systems, A/B testing and scientific experimentation, coding interview feedback, cloud-native tools, data-driven metrics
Preferred skills
Kubernetes, Ray, Kubeflow, MLFlow, gRPC, Thrift, Statsig, Meta Deltoid, Optimizely, Docker
Technologies
Kubernetes, Ray, Kubeflow, MLFlow, PostgreSQL, MySQL, DynamoDB, Apache Spark, Apache Flink, Apache Kafka, AWS, GCP, Azure, Python, C++, Go, TypeScript, gRPC, Thrift, Statsig, Meta Deltoid, Optimizely, Docker
Responsibilities
Architect scalable backend infrastructure for distributed training and inference; Lead technical design discussions and mentor engineers; Design core backend services for efficiency and low latency; Drive infrastructure optimization for compute cost and network performance; Collaborate with ML and product teams to translate requirements into robust solutions; Evaluate and integrate cloud-native technologies to enhance platform reliability
Seniority
Senior, hands-on IC with mentorship responsibilities