Job Posting Title AI/ DevOps Engineer
Core
Senior Site Reliability Engineer ensuring reliability, scalability, and operational excellence for Adobe's Real-Time Customer Data Platform (RTCDP) datastores and emerging AI/ML services.
Role type
Senior SRE (hands-on IC)
Builds
RTCDP distributed datastore ecosystem and AI/ML operational workflows for global brands
Domain
Enterprise SaaS / Data Infrastructure / AI/ML Ops
Deliverable
production ML models | infrastructure
Required skills
SRE, distributed systems, datastores (Aerospike, FoundationDB, Postgres, CosmosDB/DynamoDB), Kubernetes, cloud platforms (AWS/Azure/GCP), observability (Prometheus, Grafana, OpenTelemetry), incident response, automation, cost optimization
Preferred skills
AI/ML systems, MLOps, AI-assisted coding tools
Technologies
Aerospike, FoundationDB, Postgres, CosmosDB, DynamoDB, Kubernetes, AWS, Azure, GCP, Prometheus, Grafana, OpenTelemetry
Responsibilities
Own day-to-day reliability for RTCDP services including availability, performance, and durability; Participate in on-call rotations and incident response; Drive reliability, scaling, and operational excellence across distributed datastore platforms; Build automation-first solutions for provisioning, scaling, and lifecycle management; Support infrastructure and operational needs for AI/ML-powered services including model serving and data pipelines
Seniority
Senior, hands-on IC
