Software Engineer, Infrastructure Reliability
Core
Design, build, and operate reliable, performant, and secure distributed systems and cloud infrastructure to support global-scale AI research and products.
Role type
Senior IC Infrastructure Reliability Engineer
Builds
Core Distributed Systems, Databases, Observability, and Cloud Infrastructure platforms
Domain
AI/ML Infrastructure, Cloud Computing, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
distributed systems design, Kubernetes orchestration, cloud infrastructure (AWS/GCP/Azure), Infrastructure as Code (Terraform), observability stack management, microservices architecture, Linux administration, performance optimization, incident response, automation scripting
Preferred skills
service mesh technologies, database technologies, networking expertise, security best practices in cloud environments
Technologies
Kubernetes, Terraform, Datadog, Prometheus, Grafana, Splunk, ELK stack, AWS, GCP, Azure
Responsibilities
Design and operate reliable systems used across engineering; identify and fix performance bottlenecks to enable scaling; resolve complex technical issues; improve automation and internal tooling; contribute to incident response and postmortems
Seniority
Senior, hands-on IC with leadership experience
