Senior DevOps Engineer
Core
Design, reliability, and evolution of production infrastructure for an AI agent platform, ensuring security, scalability, and fast shipping.
Role type
Senior DevOps/SRE technical leader
Builds
Cloud-agnostic environments (Kubernetes), CI/CD pipelines, observability stacks, and internal developer tooling
Domain
Cloud infrastructure, AI/LLM orchestration, Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Linux, networking fundamentals, containerization (Docker), orchestration (Kubernetes), CI/CD, infrastructure-as-code (Terraform/Pulumi/CloudFormation), scripting (Python/Bash/Go), observability (Prometheus/Grafana/ELK/OpenTelemetry), cloud providers (AWS/GCP/Azure)
Preferred skills
service mesh (Istio/Linkerd), security & compliance (GDPR/SOC 2), GPU clusters, MLOps/LLM serving infrastructure
Responsibilities
Architect core infrastructure (networking, compute, storage), build cloud-agnostic environments, implement CI/CD, define reliability standards (SLOs/SLIs), establish observability, drive infrastructure-as-code, own security fundamentals, optimize costs/capacity, mentor DevOps team
Seniority
Senior, hands-on IC with leadership responsibilities