Senior DevOps Engineer, AI Platform
Core
Build and maintain the AI runtime infrastructure for customer-facing accounting products (Transform, AI Matching, AutoBuilder) on AWS, ensuring multi-region availability, cost control, and reliability for foundation model calls and AI-generated code execution.
Role type
Senior DevOps Engineer (AI Infrastructure)
Builds
Multi-region AI runtime, sandboxed code execution environments, and CI/CD pipelines for model and prompt deployments.
Domain
Fintech / Accounting / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS (ECS/Fargate, Lambda, SQS, S3, IAM, VPC), Terraform, CI/CD (GitHub Actions), Python/TypeScript, Observability (Grafana, Prometheus, OpenTelemetry), LLM serving (Bedrock, AgentCore, vLLM, Triton), FinOps, Security compliance (SOC 2, ISO 27001)
Preferred skills
Multi-region infrastructure under data-residency constraints, AI-specific cost optimization (token accounting, prompt caching), Sandboxed execution of untrusted code, Monorepo build systems (NX, Turborepo), Progressive delivery (feature flags), FinOps tooling (CloudZero), Data infrastructure (MongoDB, PostgreSQL, Snowflake)
Technologies
AWS Bedrock, Bedrock AgentCore, TrueFoundry, Terraform, GitHub Actions, Docker, Grafana, Prometheus, OpenTelemetry, NX, vLLM, Triton, Ray Serve, KServe, SageMaker
Responsibilities
Own AWS Bedrock and AgentCore footprint across US, EU, and AU regions; Operate sandboxed execution environments for AI-generated code; Write and review Terraform for multi-account, multi-region AWS estate; Extend observability platforms with AI-specific signals (token consumption, latency, throttle rates); Engineer cost attribution for model inference and sandbox compute; Build CI/CD pipelines for model and prompt changes; Manage on-call rotation for AI-specific failure modes; Enforce security, compliance, and tenant isolation for the AI stack
Seniority
Senior, hands-on IC