Site Reliability Engineer | AI Infrastructure
Core
Design and maintain deployment, observability, and production reliability for AI agents and LLM-backed services within a commercial real estate enterprise.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
Production AI agents for internal governance and engineering support, integrated with enterprise systems.
Domain
Commercial Real Estate / AI Infrastructure
Deliverable
production ML models
Required skills
SRE/Platform Engineering, Cloud Platforms (Azure/AWS), Containerization (Docker/Kubernetes), CI/CD, Infrastructure-as-Code (Terraform/CDK), Monitoring/Observability, Linux/Networking, Incident Management
Preferred skills
AI/ML Infrastructure (Model Serving/LLM APIs), TypeScript/Python Production Coding, Self-service Developer Tooling, Cloud Cost Optimization, Security Engineering
Technologies
Azure, AWS, Docker, Kubernetes, Terraform, CDK, CloudFormation, Datadog, Splunk, CloudWatch, TypeScript, Python
Responsibilities
Design and maintain deployment and release infrastructure for AI agents; Build monitoring and observability for AI services to track quality and costs; Implement security and compliance standards for sensitive AI workflows; Create tooling to simplify building, testing, and deploying AI services.
Seniority
Senior, hands-on IC