CareerPlanSign in

Principal Software Engineer - DevOps / Site Reliability Engineer

Singapore💼 Full-time🗓 2026-08-12 → 2026-09-26

Core

Own and evolve operational foundations for the AI Efficiency team's tech platform and tools to ensure reliability, scalability, and security at growing scale.

Role type

Principal DevOps / Site Reliability Engineer

Builds

Production AI services, internal tools, and the underlying infrastructure supporting them

Domain

Gaming industry, Cloud Infrastructure, AI Engineering

Deliverable

production ML models | infrastructure

Required skills

Cloud platform design (AWS/GCP/Azure), CI/CD pipeline design, Observability (metrics/logs/traces), Infrastructure as Code, Capacity planning, Incident management, Automation scripting (Python/Go/JS/TS), Disaster recovery design, Service resilience patterns

Preferred skills

AI-native workflow operationalization, Agentic system guardrails, Progressive delivery strategies, Self-service platform development

Technologies

Python, Go, JavaScript, TypeScript, AWS, GCP, Azure, Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK, Jenkins, GitHub Actions

Responsibilities

Design and maintain infrastructure for production AI services; Improve CI/CD pipelines and deployment automation; Define and operationalize SLOs, SLIs, and error budgets; Build comprehensive observability using metrics, logs, and traces; Lead incident response and post-mortem reviews; Automate operational toil and reduce manual intervention; Design resilience and disaster recovery strategies; Improve developer experience via self-service tooling; Partner with ML engineers on model-serving reliability; Implement AI-assisted operational workflows with proper guardrails; Mentor team members on operational excellence.

Seniority

Principal, hands-on IC with strategic leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.