Sr. Member of Technical Staff
Core
Design and develop software features for system resiliency, high availability, and scalable AI inference services in distributed environments.
Role type
Sr. Member of Technical Staff (Infrastructure & Reliability)
Builds
Cloud-based deployment workflows, automated recovery mechanisms, and high-performance inference software.
Domain
AI Infrastructure / Cloud Systems
Deliverable
production ML models | infrastructure
Required skills
Python, Terraform, Kubernetes, AWS services, Docker, CI/CD pipelines, distributed tracing, fault-tolerant architecture
Preferred skills
Node.js, PostgreSQL, Redis, Grafana, Helm, asynchronous processing
Technologies
AWS (EC2, Lambda, EKS, ECS, Fargate, CloudWatch, X-Ray), Docker, Kubernetes, Terraform, Ansible, Jenkins, Git, Python, Node.js, Flask, PostgreSQL, Redis, ELK stack, Prometheus, Grafana
Responsibilities
Design fault-tolerant architecture for distributed AI inference; develop cloud deployment workflows using AWS; create Python scripts for data preprocessing and inference execution; implement automated recovery and monitoring tools; debug deployment and networking issues; document technical configurations and workflows.
Seniority
Senior, hands-on IC